EuraStudy
EuraStudy assistant
Skip to content
EuraStudyThe Lab
CurriculaMethodsAreasThe workFrom the field
EuraStudy →

EuraStudyResearch & engineering

Built on
evidence.

The ideas behind EuraStudy. Learning science, AI research and the engineering that brings them together.

Explore the work ↓Our methods ↗︎
Research & engineering
Research & engineering↓

The open notebook Latest entry · 8 September 2026

15
Entries
4
Areas
8
Cited works
60
Computed figures

§ 01 · Methods

Methods we build on.

The instruments the platform actually runs on — the evidence behind each is cited under From the field.

HIGHLOWUNAIDED PERFORMANCEPRACTICE OVER TIME →DELAYED POST-TESTANSWER-GIVING TUTORfast feeling of fluencySCAFFOLDED RESTRAINTslower, then durableproductive difficultyCROSSOVERwhere struggle pays offFig. mastery over practice

Knowledge tracing

A per-topic estimate of what a student has really mastered, updated with every answer.

Read · How a Machine Reads What You Know→
CAN DOALONECAN DO WITH HELP— the zone of proximal development —CANNOT YETGOOD TUTORING OPERATES HEREMASTERYexpands outwardFig. the zone of practice

Adaptive testing

Every question carries a calibrated difficulty; the diagnostic picks the item that reveals the most.

Read · Twenty Questions→
RETENTION — REVIEW BEFORE YOU FORGET100%0FIRST STUDYTIME →REVIEW THRESHOLDR1R2R3R4each review, a flatter forgettingFig. review before you forget

Spaced repetition

A review timed just before forgetting lifts memory back and flattens the next decay.

Read · The Half-Life of Knowing→

§ 02 · Areas

Four areas of work.

A·0102AI tutoringHow the tutor decides what to say — and, more often, what to withhold.2 entries · latest: The Grammar of a Hint→A·0205Assessment & feedbackMeasuring answers, marking like an examiner, and judging the tutor itself.5 entries · latest: When the Tutor Leaves→A·0303Learning scienceWhat the evidence on practice, memory and difficulty actually supports.3 entries · latest: Reliability Is a Promise About Noise→A·0405Design & craftThe instruments, typography and figures the whole platform is built from.5 entries · latest: Every Millisecond Is a Promise→

§ 03 · The work

The work.

15 of 15
ResearchAssessment & feedbackWhen the Tutor LeavesWhat does an 80% practice score mean when AI helped produce it? An original analysis of 2,899 student sessions examines the distance between assisted success and unaided performance.8 September 2026 · 16 min readRead →

When high practice scores meet an unaided exam

Low unaided exam scores after high practice scoresSchematic diagram with 24 elements, 25.6%, Control, 22 / 86, 64.5%, GPT Base, 80 / 124, 66.7%, GPT Tutor, 318 / 477, 0%, 25%, 50%, 75%, 100%25.6%Control22 / 8664.5%GPT Base80 / 12466.7%GPT Tutor318 / 4770%25%50%75%100%
Low unaided exam scores after high practice scoresSchematic diagram with 21 elements, 25.6%, Control · 22 / 86, 64.5%, GPT Base · 80 / 124, 66.7%, GPT Tutor · 318 / 477, 0%, 25%, 50%, 75%, 100%25.6%Control · 22 / 8664.5%GPT Base · 80 / 12466.7%GPT Tutor · 318 / 4770%25%50%75%100%

Exam below 50% among sessions with practice ≥80%. Whiskers: 95% classroom-bootstrap intervals.

Fig. D·01
ResearchAI tutoringThe Grammar of a HintA good hint is a sentence with a very particular job: to make the next move the student’s own. Too little and they stay stuck; too much and the problem is solved for them, which is to say not solved at all. We read the tutoring literature for the grammar of help — zones, rungs, fading, the assistance dilemma — and describe the ladder our tutor climbs, one rung at a time, only when asked.Read →
PLATE · THE HINT LADDERASSISTANCE DILEMMA0.90.50.2FINISH UNAIDED · PSOLUTION REVEALED →01 · REORIENT0%02 · NARROW25%03 · STEP60%04 · FULL PATH100%a reorientation costs almostnothing — and teaches almost nothingthe full path solves the problem —and ends the practiceHELP · SCAFFOLDED IN RUNGSP(finish) = 0.92 − 0.72·reveal
Fig. D·02
22 Aug · 9 min read
ResearchAssessment & feedbackWhat a Wrong Answer Is WorthEvery exam paper returns a number; almost none of them returns a plan. A wrong answer is information about what to do next — arguably the most actionable information a student ever receives — and most systems throw it away the moment the mark is recorded. We follow the error from verdict to diagnosis to scheduled repair, and read the evidence on feedback, failed attempts and productive struggle for what a mistake is actually worth.Read →
PLATE · THE CORRECTION LEDGEREXPANDING REPAIR LADDER1 d3 d7 d14 d14 dATTEMPTthe error happensD0D1REPAIRre-attempt · widenedD4D11REPAIRD25SETTLEDD39VERDICT · DIAGNOSIS · SCHEDULEan error enters the ledger; the ledger decides when it is asked againNOT A GRADE · A QUEUEGAPS WIDEN AS REPAIRS HOLD
Fig. D·03
18 Aug · 8 min read
EngineeringDesign & craftEvery Millisecond Is a PromiseLatency is not a performance metric; it is a design material, and every millisecond spent is a sentence spoken to the student. A hover that waits half a second says the interface did not notice. A tutor that answers in two hundred milliseconds says listening is optional. We read the human–computer literature for what each scale of waiting means, and the classroom literature for why the best teachers pause — then draw our own budget for the first second of a signed-in page.Read →
PLATE · THREE LANES OF WAITINGMILLER · 196810 ms50 ms1 s5 s10 sRESPONSE TIME · LOG SCALEINSTANTthe hover light · a chip · a togglerespond in placeCONVERSATIONa streamed tutor token · search as you typeacknowledge at once, finish when readyTHE THREAD BREAKSa graded essay · a built report · a generated examshow progress, keep the promise visibleBEYOND WORKING MEMORYLATENCY IS A DESIGN MATERIALEVERY MILLISECOND SPENT IS A PROMISE MADE
Fig. D·04
14 Aug · 7 min read
ResearchLearning scienceReliability Is a Promise About NoiseEvery test score is two numbers pretending to be one: the student, and the noise. Measurement theory is the discipline that keeps the two apart — Spearman’s decomposition, Cronbach’s ratio, the standard error that turns a point into a range, and Kane’s argument chain that turns a range into a decision you can defend. This dispatch is why our diagnostic reports bands instead of points, and why an honest instrument would rather say “between” than guess.Read →
PLATE · THE ESTIMATE AND ITS RANGE± 1.96 · SEM = SD / √N-20+210 ITEMSa coin-flip of placement-20+220 ITEMSa shorter range-20+245 ITEMSa claim you can act oncertainty is bought by length — twice the precision, four times the questionsCLASSICAL TEST THEORYBANDS COMPUTED FROM THE √N LAW
Fig. D·05
10 Aug · 8 min read
EngineeringDesign & craftA Calculus of DiagramsMost diagrams in educational software are pictures someone drew once. Ours are values in a typed language, compiled to pixels by a function that cannot lie. A formal account of the figure engine — its grammar, its determinism, and the proof obligations that keep nearly two thousand diagrams honest.Read →
THE PIPELINE — A CHAIN OF TOTAL FUNCTIONSPLATE I — render : Spec → SVGREJECTunknown kind · malformedREJECTunsafe markup01SPEC02VALIDATEa gate03RENDER04SANITISEa gate05STORE06DISPLAYevery arrow either succeeds with a typed value, or is refused at a gate
Fig. D·06
20 Jun · 9 min read
ResearchLearning scienceThe Half-Life of KnowingYou can know something on Tuesday and not know it on Friday — memory has a half-life, and it is shorter than anyone would like. Spaced repetition is the engineering discipline built on that uncomfortable fact: schedule each review for the moment a memory is about to fade, and a little forgetting becomes the thing that makes learning stick. We trace the idea from Ebbinghaus to the algorithms now built into the tools millions revise with.Read →
PLATE · THE FORGETTING CURVEEXPONENTIAL · S 6 d1.00.50.0R · RETRIEVABILITY051015202530DAYS SINCE STUDY →HALF-LIFE ≈ 4.2 dS · STABILITY = 6 drecall has fallen to 1/efreshly learned · R = 1without review, even well-learnedmaterial leaks awayMEMORY · LEFT ALONER(t) = e^(−t / S)
Fig. D·07
19 Jun · 14 min read
ResearchAssessment & feedbackTwenty QuestionsA good adaptive test can pin down what you know in a dozen questions, not fifty — because it chooses each one to be the most revealing it can ask. We trace the quiet mathematics of item response theory and computerized adaptive testing, from the shape of a single question to the loop that learns you in real time, and the places where adaptivity has to be reined in.Read →
PLATE · THE ITEM CHARACTERISTIC CURVE3PL · a 1.3 · b 0.4 · c 0.201.00.50.0P( CORRECT )-3-2-10+1+2+3ABILITY θ →c · GUESSING FLOOR = 0.20even a pure guess sometimes landsb · DIFFICULTYa · DISCRIMINATIONthe slope through the midpointmidpoint · P = (1 + c)/2ONE ITEM · ONE CURVEP(θ) = c + (1−c)·σ( a(θ−b) )
Fig. D·08
18 Jun · 13 min read
ResearchAssessment & feedbackHow a Machine Reads What You KnowEvery adaptive tutor rests on a quiet act of inference: guessing the knowledge it cannot see from the answers it can. We trace that idea from Bayesian Knowledge Tracing to its deep-learning successors — and the honest places where the deeper model is not the better one.Read →
THE MODEL · A TWO-STATE HIDDEN MARKOV MODELHIDDEN STATE — NEVER OBSERVEDOBSERVED — THE ONLY THING THE TUTOR SEESp(L0) · PRIORNOT KNOWNKNOWNp(T) · LEARNno forgetting1−p(T)INCORRECTCORRECT1−p(G)1−p(S)p(S) · SLIPp(G) · GUESS
Fig. D·09
17 Jun · 9 min read
EngineeringDesign & craftOn the Making of a Quiet MachineA study of the obsessions behind a learning platform built for four national examinations — where nothing is accidental, and restraint is the most exacting discipline of all.Read →
THE SINGLE ACTIONTHE STAT LEDGERAaTHE SERIF NAMEF·07THE INDEX CODEPLATE I — ANATOMYEXPLODED · 1:1
Fig. D·10
14 Jun · 7 min read
ResearchAI tutoringWithholding the AnswerA system that hands over the answer is not teaching. We argue that the central design problem for a machine tutor is not how to explain, but when and how much to withhold.Read →
LEARNERLTUTORTQUERY · ATTEMPTFEEDBACK · HINTt₀ · FULL SUPPORTtₙ · INDEPENDENCE
Fig. D·11
11 Jun · 9 min read
EngineeringDesign & craftDrawn, Not DecoratedEvery chart, curve and diagram a student meets is drawn to exact specification by a single figure engine — and verified before it ships. Never faked, never screenshotted.Read →
VERIFICATION — EVERY QUESTION, BEFORE A STUDENTA QUESTION01DRAFTEDauthored02SOLVEDre-worked03ADJUDICATEDreviewed04CRITERIA SUMpoints tally05PUBLISHEDmade visibleSEALED
Fig. D·12
8 Jun · 5 min read
ResearchAssessment & feedbackHow Should We Measure a Tutor?A tutor that keeps students busy is not the same as a tutor that helps them learn. We argue for measuring AI tutors by learning gains and transfer — and against the engagement metrics that quietly reward the wrong thing.Read →
HIGHLOWUNAIDED PERFORMANCEPRACTICE OVER TIME →DELAYED POST-TESTANSWER-GIVING TUTORfast feeling of fluencySCAFFOLDED RESTRAINTslower, then durableproductive difficultyCROSSOVERwhere struggle pays off
Fig. D·13
5 Jun · 8 min read
EngineeringDesign & craftOne Platform, Four National ExamsThe Austrian Matura and the German Abitur are live; the French Baccalauréat and Spanish Selectividad are on the waitlist. The hard part was never the content — it was deciding what four exams could share without flattening any of them.Read →
THE ATLASFour Examinations, One SkyÖSTERREICHMATURADEUTSCHLANDABITURFRANCEBACCALAURÉATESPAÑASELECTIVIDAD
Fig. D·14
2 Jun · 6 min read
ResearchLearning scienceAdaptive Practice and Its LimitsAdaptivity is the most over-promised word in educational technology. Two effects in the learning-science record are real and worth building on; almost everything sold above them is decoration.Read →
RETENTION — REVIEW BEFORE YOU FORGET100%0FIRST STUDYTIME →REVIEW THRESHOLDR1R2R3R4each review, a flatter forgetting
Fig. D·15
28 May · 8 min read

Selected research · 8 works

From the field.

A standing reading of the research on artificial intelligence and learning — the work of others, across decades, that the rest of this notebook is built on. These are published findings by researchers across the field, not EuraStudy’s own results; we summarise them and point to the original work.

  1. 01Tutoring efficacy

    One-to-one tutoring moved the average student to the 98th percentile.

    Students who worked with a personal tutor outperformed conventionally taught peers by about two standard deviations — Bloom’s “two sigma” result. It set the central ambition that has driven educational technology ever since: to reproduce, at scale, what a good tutor does for one learner.

    Benjamin S. Bloom · 1984 · The 2 Sigma Problem · Educational Researcher

  2. 02Tutoring efficacy

    Intelligent tutoring systems came within a whisker of human tutors.

    Reviewing decades of controlled studies, VanLehn measured human tutoring at roughly 0.79 standard deviations over no tutoring and step-based intelligent tutors at about 0.76 — far below Bloom’s famous 2.0, and close enough to each other to reframe the question from “can a machine tutor?” to “what does effective tutoring actually consist of?”

    Kurt VanLehn · 2011 · The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems · Educational Psychologist

  3. 03Evidence & meta-analysis

    Across fifty controlled evaluations, intelligent tutors raised scores by about two-thirds of a standard deviation.

    The median system raised scores by about two-thirds of a standard deviation — but far more on the locally designed tests that match what a system actually taught (around 0.73) than on standardised exams (around 0.13). Real, and a reminder that the size of an effect depends heavily on what you choose to measure.

    James A. Kulik & J. D. Fletcher · 2016 · Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review · Review of Educational Research

  4. 04Memory & practice

    Being tested on material beats re-reading it — and the gap widens with time.

    Learners who practised retrieving what they had studied remembered substantially more a week later than those who simply restudied — even though the restudiers felt more confident at the time. The “testing effect” is among the most robust results in the science of learning, and the reason deliberate practice, not mere exposure, sits at the centre of exam preparation.

    Henry L. Roediger III & Jeffrey D. Karpicke · 2006 · Test-Enhanced Learning · Psychological Science

  5. 05Cognitive load

    Working memory is the bottleneck — and the help a novice needs becomes noise for an expert.

    Cognitive load theory holds that instruction fails when it overwhelms a narrow working memory. Later work on the “expertise-reversal effect” sharpened the point: scaffolding that helps a beginner actively hinders a more advanced learner. Together they argue that good tutoring must adapt its support to the individual, not just to the topic.

    John Sweller · 1988 · Cognitive Load During Problem Solving: Effects on Learning · Cognitive Science

  6. 06Learning theory

    Good help is temporary: a scaffold exists in order to be removed.

    Wood, Bruner and Ross named “scaffolding” — the support an expert lends so a learner can do what they cannot yet do alone, an idea since drawn together with Vygotsky’s zone of proximal development. Its defining feature is that it fades: support that never withdraws breeds dependence, not competence. It is the principle behind any tutor that deliberately holds back the answer.

    David Wood, Jerome S. Bruner & Gail Ross · 1976 · The Role of Tutoring in Problem Solving · Journal of Child Psychology and Psychiatry

  7. 07Feedback

    Feedback is one of the most powerful influences on learning — and one of the most variable.

    Synthesising hundreds of studies, Hattie and Timperley placed feedback among the strongest levers on achievement, with effects ranging from large to outright negative. What separated them was whether the feedback told a learner where they were going, how they were doing, and what to do next. Feedback that grades without directing can achieve nothing at all.

    John Hattie & Helen Timperley · 2007 · The Power of Feedback · Review of Educational Research

  8. 08Critical perspectives

    Perhaps the machine should stay simple, and the intelligence should stay human.

    Baker argues the field over-invested in modelling the learner’s mind and under-invested in the simpler, robust systems that actually help — and in keeping teachers in the loop. A standing corrective for anyone building an AI tutor: sophistication is not the goal; better learning is.

    Ryan S. Baker · 2016 · Stupid Tutoring Systems, Intelligent Humans · International Journal of Artificial Intelligence in Education

Selected reading · 8 works · a starting point, not a survey

Our editorial standard

Evidence you
can follow.

A useful idea should stand up to a closer look. Here is how we make the work in this notebook traceable.

01 / Figures

Drawn from the method.

Computed to specification and verified before publication.

How we build figures ↗︎
02 / Sources

Connected to the source.

Published work, named researchers, and references you can follow.

Explore the reading list ↑
03 / Limits

Clear about uncertainty.

What the evidence supports, where it stops, and what remains open.

Read about the limits ↗︎
EuraStudy

Bringing AI into Europe's education, from final exams to university.

Contact supporteurastudy@gmail.com

Products

EuraStudyThe exam studio for twelve European school-leaving exams.
  • Features
  • Gubernik

Exams

  • Matura · Österreich
  • Abitur · Deutschland
  • Bac · France
  • Selectividad · España
  • Maturità · Italia
  • Exames Nacionais · Portugal
  • A-Levels · UK
  • Leaving Cert · Ireland
  • Matura · Polska
  • Πανελλαδικές · Ελλάδα
  • Havo · Nederland
  • Vwo · Nederland

Company

  • About
  • Research
  • News
  • Guides
  • FAQ
  • Contact

Legal

  • Privacy policy
  • Terms of use
© 2026 EuraStudy·All rights reserved.Made in Austria, for Europe