The science

Harder in the right places. Easier everywhere else.

Workings isn’t built on ed-tech fashion. Each rule in the engine comes from the cognitive science of learning, and each one is written into the code rather than left as a setting. Here is the science, in plain language, and exactly what it makes the product do.

Uncle
Nice try! I’ll never just tell you — but I’ll help you get there.
Kit
You said 33. Have a look at the ones column with me. What is 2 take away 5?
Marlow
You’ve got Year 3 subtraction. That’s yours now.

Productive struggle · the generation effect

Never the answer. Always the scaffold.

The science

Memory is built by retrieving and generating, not by reading someone else’s solution. Answers you produce yourself are remembered better than answers you are shown — the generation effect. And curiosity is not just a mood: it arises from a gap that looks closable, and that state helps the brain encode what comes next.

Handing a child the answer closes the gap and switches off the very state that was helping them learn.

What Workings does

  • The answer is never revealed, in any mode. That was decided, and it is enforced by software checks, not by good intentions.
  • When a child is stuck, help comes as a ladder of smaller steps, and the child always takes the last one.
  • Every question and hint is checked so it can never contain the answer, and the answer is never sent to the child’s browser.
  • If a child disengages, the question closes unanswered. It isn’t counted, and its maths returns another day in a new form.

The hint ladder

  1. A question aimed at the mistake

    If a wrong answer matches a known misconception, Kit asks the one question that exposes it. Otherwise Pip prompts: what do you already know?

  2. A hint, not a step

    Still stuck? A hint at the next tier — shorter as the tiers rise, and never containing the answer.

  3. If the reading is the snag

    When the engine puts the difficulty down to reading, Bramble unpacks the question instead of giving a maths hint.

  4. Step back to the gap

    Deeper still, a smaller question from the prerequisite the child needs — then straight back to the original.

  5. The last step is always theirs

    Fishing for the answer earns Uncle’s “nice try”. If a child walks away, the question closes unanswered and its maths comes back another day in a new form.

Desirable difficulties · error feedback

The child must sometimes be wrong.

The science

Learning runs on prediction error: no discrepancy, no signal, no update. Practice that feels easy often produces less lasting learning than practice with the right amount of difficulty — what Elizabeth and Robert Bjork call desirable difficulties.

Research on how learners improve suggests the most efficient training sits at roughly 85% success, or about 15% errors (Wilson and colleagues, 2019). Workings allows extra headroom for young and anxious learners.

What Workings does

  • Before serving a question, the engine predicts this child’s chance of success and picks one in the 70–85% band, aiming for 78%.
  • Ability and question difficulty are both estimated from every answer, in the style of Elo ratings used by adaptive systems such as Math Garden.
  • A session where a child gets everything right is treated as mis-pitched. Sustained success above 90% pushes the difficulty up.
  • The session opens with confident warm-ups and always closes on a question the engine expects the child to get right.

Predicted chance of getting it right

Every question is chosen to land between 70% and 85%, aiming for 78%. Below it, practice hurts. Above it, nothing new is being learned.

Cognitive load theory · working memory

One stretch at a time.

The science

Working memory holds only a handful of items at once — about four chunks. Whatever competes for those slots doesn’t get learned. If most of them are spent decoding a sentence, the arithmetic has nowhere to run. That’s cognitive load theory in a line.

Vocabulary adds load too. Researchers sort words into everyday words, general academic words and subject-specific words (Beck, McKeown and Kucan), and the last two cost a young reader far more.

What Workings does

  • Every question sits on a five-level reading scale, from no reading at all to inference, with measured limits on words, vocabulary and distractors.
  • A question may push the maths or the reading, never both at once.
  • Each child has two abilities — maths, and reading in maths — so a wrong answer is attributed to whichever was more likely the cause.
  • Hints are never harder to read than the child is comfortable with, and read-aloud never counts against them.

The same maths, five reading loads

  1. Reading level 1: No reading

    72 − 45 = ?

  2. Reading level 2: One plain idea

    Jarrah had 72 cards. She gave away 45. How many are left?

  3. Reading level 3: A short situation

    Jarrah had 72 footy cards. She gave 45 of them to her cousin. How many cards does Jarrah have left?

  4. Reading level 4: A situation with noise

    Jarrah’s team has 18 players. She had 72 footy cards and gave 45 to her cousin. How many does she have now?

  5. Reading level 5: Inference

    Jarrah now has 45 fewer footy cards than at the start of the season, when she had 72. How many does she have?

Every line asks for 72 − 45. The numbers and the operation never move; only the reading does.

Retrieval practice · the spacing effect · consolidation

Remembered next month, not just today.

The science

Pulling something out of memory strengthens it more than looking at it again (retrieval practice), and spreading that practice over widening gaps beats cramming (the spacing effect). Both are among the most reliable findings in learning research (Dunlosky and colleagues, 2013).

Sleep matters as well: it’s part of what moves a skill from effortful to automatic. Performance during practice is not the same as learning.

What Workings does

  • Mastered skills come back after 1, 3, 7, 16 and 35 days, at the harder end of the stage and in a different question type where possible.
  • The same maths returns in new forms; the identical question never repeats for a child, ever.
  • Crossing the mastery line isn’t enough: it must be confirmed by a correct, unhinted answer after at least one night’s sleep.
  • Sessions are short — about 20 minutes — and then they stop.
  1. Mastered
  2. Day 1
  3. Day 3
  4. Day 7
  5. Day 16
  6. Day 35
A mastered skill comes back at widening intervals, in a different form each time. Two missed reviews in a row put it back in practice — never shown as a loss.

Interleaving

Mixed practice, not blocks of the same thing.

The science

Doing twenty of the same problem type in a row feels productive, but mixing problem types forces a learner to decide which method a problem needs — and that choice is much of what maths asks. Interleaved maths practice has outperformed blocked practice in classroom studies (Rohrer and Taylor, 2007).

What Workings does

  • Practice is interleaved across strands, with no more than three questions in a row from one skill.
  • Roughly 70% of questions are at the child’s frontier, about 20% are spaced review, and remediation steps in when a child stalls.
  • Question types rotate too, so the same idea is met as a picture, a number sentence, a word problem and a riddle.

Concrete, pictorial, abstract

Many ways to see the same number.

Children build understanding by moving between hands-on objects, pictures and symbols — the Concrete–Pictorial–Abstract progression that grows out of Jerome Bruner’s work. Workings asks the same maths through ten question types, from hands-on to explain-it, and draws every model precisely from structured components, never with an image model.

Base-ten blocks72 as tens and ones. Take 45 away — what’s left?
7245?
Bar modelA whole of 72, a part of 45, and a part to find.
45507072+5+20+2
Number lineJump from 45 to 72. How far did you go?

Error analysis · metacognition

Every mistake is a clue.

The science

Children’s wrong answers are rarely random. 72 − 45 = 33 is what you get by taking the smaller digit from the larger in each column — a specific, fixable misconception. Asking children to find and explain errors, and to justify their own thinking, builds the reasoning that syllabus writers call working mathematically and that Webb’s Depth of Knowledge rates as higher-order.

What Workings does

  • The wrong options in multiple-choice and spot-the-mistake questions are real errors, each mapped to a misconception.
  • When a child’s answer matches one, Kit asks the one question that exposes it.
  • “Spot the mistake”, “riddle” and “explain it” questions sit alongside the procedural ones, and reasoning types count for more in mastery.
  • Each week, common wrong answers no one anticipated are proposed as new misconceptions, checked against children’s real answers, and added only when a teacher approves.

Spot the mistake

Tom says 72 − 45 = 33. Is he right? What did he do?

Kit
You said 33. Have a look at the ones column with me. What is 2 take away 5?

Kit’s line is from an illustration of a session. The answer is never given.

Prerequisite knowledge · remediation

A curriculum that knows what depends on what.

The science

New knowledge is built on old knowledge. When a child is stuck on a topic, the obstacle is often a prerequisite underneath it — and more practice on the topic itself won’t fix that.

Fractions are a good example: in the Australian syllabus the groundwork is partitioning a length in measurement, which is why fractions are taught on a number line rather than as slices of pizza.

What Workings does

  • The curriculum is a graph of skills by year and strand, with each link marked as required or supporting.
  • A skill unlocks only when its required prerequisites are mastered, so a child progresses per strand and is never capped at their school year.
  • When a child stalls, practice drops to the weakest prerequisite for a few questions, then returns.
  • The graph, the dependencies and the mastery bar are owned by a mathematics teacher, not by the software.
A slice of the curriculum graphYear 2 Geometric measure is required for Year 3 Fractions. Year 3 Multiplicative relations supports Year 3 Fractions. Year 3 Fractions and Year 4 Multiplicative relations are both required for Year 4 Fractions.Year 2Geometric measureYear 3MultiplicativeYear 3FractionsYear 4MultiplicativeYear 4FractionsStalled here?Step back topartitioning a length.
Solid arrow: requiredDashed arrow: supportingGreen: the gap to step back to

Measuring mastery

Mastery is estimated, not counted.

The science

Because Workings deliberately serves questions a child gets right about 78% of the time, a raw hit rate says almost nothing about mastery. The better question is: if this child faced every kind of question this stage asks, how often would they succeed?

What Workings does

  • Mastery is calculated across every complexity level, every valid question type and the reading level the stage expects, weighted toward the harder end and the reasoning types.
  • It needs at least 20 responses on a skill, so a lucky run can’t cross the line.
  • Guessing is left out: multiple choice counts for less, because it’s weaker evidence.
  • The mastery threshold is set by the curriculum lead, and a mastery once earned is never taken away retrospectively.

Maths anxiety · motivation

No clock. Nothing to lose.

The science

Maths anxiety isn’t a metaphor. In anxious children the threat response fires in anticipation of a timed task, and the anxiety consumes the working memory the arithmetic needs (Ashcraft, 2002; Lyons and Beilock, 2012).

The mechanic the research warns about isn’t counting effort. It’s the count you can break.

What Workings does

  • No timers, countdowns, speed scores or “quick!”. Response time is never used as evidence of ability.
  • The 20-minute budget is kept invisibly; a child never sees a clock or a count of questions remaining.
  • Reps and streak weeks only accumulate. A week that falls short takes nothing away, so there is no run to break and no “at risk” warning.
  • Over a term, recognition tapers from counts toward mastery unlocks, so motivation moves into what a child can do.
  • No leaderboards and no comparison with other children. The words are “next” and “frontier”, never “behind”.

Verifiable by design

The AI writes the words. It never decides the answer.

The science

Language models are fluent and occasionally wrong. A model that miscounts is an inconvenience; one that marks a child wrong for a right answer is a betrayal. The only safe design is one where a model can’t decide what’s correct.

What Workings does

  • Every answer is computed by code from the question’s own numbers. Fractions and decimals use exact arithmetic.
  • Question wording is drafted offline, checked by software across fifty sets of numbers and approved by a person. No model is called while a child waits.
  • Every question is validated before it is shown — no invented numbers, no leaked answer, no reading load above its level — and anything that fails is never shown.
  • A child’s name, email and school are never sent to a model.

What we don’t claim

Honest about the evidence.

The research above is about how learning works. It isn’t proof that Workings works — that has to be measured.

  • We don’t publish efficacy percentages or “grade levels of growth”. When we have independent results, we’ll share them.
  • We plan an independent, pre-registered evaluation with consented, de-identified data, and the engine checks its own calibration continuously.
  • Workings is not a clinical product. It doesn’t diagnose, treat or remediate any condition.
  • Streaks and reps are our weakest-evidence decision, so they ship as an experiment with written criteria for removing them.
Coming soon

Practice that respects how children learn.

Fourteen days free. Twenty minutes a day. Every question pitched for your child, and the last step always theirs.