---
title: "How Kerfox works - spacing, retrieval practice and marking to the mark scheme"
description: "The engines under Kerfox: an FSRS scheduler that brings each topic back before it slips, questions pitched where you get about eight in ten right, and marking against the exam board’s own scheme."
image: "https://www.kerfox.app/og.webp"
url: "https://www.kerfox.app/science/how"
---
[Kerfox](https://www.kerfox.app/)

THE RESEARCH BEHIND KERFOX

# Why we built it this way.

We have no evidence that Kerfox works. Nobody has run a study on it, no class has been taught with it and measured against a class that was not, and until somebody does, any number we put on this page would be one we made up.

What we do have is this. Every design decision in the app was taken from a published finding, and we can show you which one. Below, each method is named, the research behind it is cited and linked, and we say exactly where in the app you run into it, so you can go and look rather than take our word for it.

That is a claim about our methods rather than about your marks. It is the only honest claim available to a study app that nobody has studied yet, and we would rather make it well than make a bigger one badly.

- 36papers read, linked and summarised
- 9methods, each traced back to them
- 43checks that fail the build rather than warn
- 0studies of our own

The last one is the number that decides how the rest of this page is written.

WHAT THIS PAGE IS NOT

It is not an efficacy page. The findings below were established by other people, in their own studies, on their own materials. That a technique works in a published experiment is not evidence that our implementation of it works, and we are careful throughout to keep those two things apart.

It is also not a sales brochure with footnotes. Where a finding is contested, or where its scope is narrower than the headline suggests, we say so in the same paragraph rather than quietly rounding it up. Three of the methods below carry a sentence that cuts against them.

The engines running underneath, the teaching methods they implement, and the things we built and then took back out.

## What is actually running.

Under the app there are a handful of small engines, and none of them is mysterious. Each one is here because a decision had to be made hundreds of times a day and making it by hand would have meant making it badly. What follows is what each one is for, what it is aiming at, and where we have deliberately not put a model at all.

This is the level of detail a teacher asked us for and the level we are comfortable at: what the thing is doing and why, without the parameters we are still tuning.

### The scheduler

Decides the day a topic comes back.

Every topic carries a small model of your memory of it - how hard you find it, how strong the memory is now, and how fast that is decaying - and the model is asked when the memory is about to slip rather than after it has. We use the published FSRS scheduler rather than intervals of our own, and that is a deliberate refusal to be clever: a hand-built ladder of days is a guess wearing the same confidence as a fitted model, with none of the fitting behind it. On top of it sits one clamp of ours, from the spacing ridgeline result: as an exam approaches, the gap between reviews is pulled toward a fraction of the time left instead of being allowed to run out past the paper.

SOURCES[Murre & Dros (2015)](https://www.kerfox.app/science/how#paper-murre2015)[Cepeda et al. (2008)](https://www.kerfox.app/science/how#paper-cepeda2008)[Ye, Su & Cao (2022)](https://www.kerfox.app/science/how#paper-ye2022)

### The pairing

Decides which question you get next.

You have an estimated ability per topic and every item has an estimated difficulty, so the chance you get a given item right can be predicted before it is served. What that prediction is aimed AT depends on why the item is being served, and the three jobs are genuinely different. A diagnostic aims at 0.60, near the point where an answer carries the most information about you but backed off from it, because a probe pitched at a coin flip measures beautifully and feels like being told you know nothing. A teaching item aims at 0.85, which is the Eighty Five Percent Rule. Retrieval near a paper aims at 0.80 and ramps down toward 0.70 in the final fortnight, because close to an exam the job is to find what will fail under pressure rather than to feel fluent.

SOURCES[Wilson et al. (2019)](https://www.kerfox.app/science/how#paper-wilson2019)[Bjork & Bjork (2011)](https://www.kerfox.app/science/how#paper-bjork2011)

### A weight, never a filter

Decides what happens when nothing fits.

The pitch above is applied as a weighting on the pool rather than as a gate on it, and that one decision is doing more work than the targeting is. A filter on a bank as young as ours produces an empty session, and an empty session is worse than a badly pitched one by a distance no amount of pedagogy closes. Weighting degrades exactly gracefully: when the bank has something ideal you get it, and when it does not you get the nearest thing rather than nothing. The honest reading of that is that our difficulty targeting is at its weakest on the topics with the fewest questions, which is the same set of topics where it would help most.

### The marker

Decides whether what you typed is right.

Almost nothing here is compared as text. A number is marked on its dimension as well as its value, so five joules is refused against five newtons rather than passed on the digit. A formula is EVALUATED - yours and ours, at a dozen scattered points - and marked equal if they agree everywhere, which is why every correct rearrangement scores and why mgh, half m v squared and root 2gh each read as the physics they are. A comparison like v is less than c is marked by where it is TRUE rather than by its shape, so writing it the other way round is the same answer. An integral is marked by DIFFERENTIATING what you wrote, which is the only way the whole family of antiderivatives passes and nothing outside it does. A derivation is marked line by line, each line tested for truth and then for which physics it actually used.

### The diagnosis

Decides what to say about a wrong answer.

A wrong answer that is a fixed multiple of the right one is not a random miss, it is a named mistake: a missing one half, degrees where radians were wanted, a reciprocal, a stray factor of two pi. Those are recognised from the ratio itself with nothing authored, so a slip gets named whether or not whoever wrote the question thought of it. Authoring is reserved for the misconceptions only the author knows about, and none of it ever changes a mark.

### The mark scheme

Decides what a score means.

Answers are not marked right or wrong, they are marked OUT OF MARKS, the way the boards do it. That is why the border of an answer box is cut into one arc per mark and the arcs land one at a time: three out of four is a fact you can read off the edge without reading a number. At GCSE and A-Level it goes further, because at those levels an answer is credited for particular WORDS - denatured scores where killed does not, though both describe an enzyme that has stopped working - so the lessons hand over the specification's own wording at the moment the idea it names is finished being built.

SOURCES[Hattie & Timperley (2007)](https://www.kerfox.app/science/how#paper-hattie2007)[Shute (2008)](https://www.kerfox.app/science/how#paper-shute2008)[Butler, Godbole & Marsh (2013)](https://www.kerfox.app/science/how#paper-butler2013)

### Where there is no model at all

The part most people assume is a model, and is not.

Nothing you type is sent anywhere to be marked. Every mark in this app is decided by arithmetic running on your own device, which is why it works on a plane, on a dead signal and in an exam hall. A language model is used in exactly two places and you have to ask for both: turning notes you paste in into practice, and adding a note about the points a written explanation missed. Both fall back to a local result the moment they are slow or unavailable, on a six second limit, because a mark that waits on a request which may never arrive is not a mark. We are not claiming this is more advanced. It is less advanced on purpose. Where those two features DO send something, our terms say what happens to it - including that notes pasted into the AI generator go to Google and may be kept - and that page is the one to read before you paste anything you would mind being kept.

## The methods, and where you meet them.

Each of these is the same three things in the same order: what the research found, what we built because of it, and where it turns up while you are working. The third one is the only part you can check for yourself, so it is the one we have tried hardest to make specific.

### Spacing, and the shape of forgetting

#### WHAT THE RESEARCH FOUND

Memory for something learned once falls away steeply in the first hours and then much more gently, a curve first measured in the 1880s and replicated closely in 2015. Spreading study sessions apart rather than stacking them together is one of the most reliably reproduced findings in the field: across hundreds of experiments, spaced learners scored 47 per cent on the final test against 37 per cent for massed ones. How far apart is the interesting part. The best gap depends on how long you need to hold the material, and it is a shrinking fraction of that delay: roughly a fifth of it when the test is weeks away, nearer a twentieth when it is a year away.

SOURCES[Murre & Dros (2015)](https://www.kerfox.app/science/how#paper-murre2015)[Cepeda et al. (2006)](https://www.kerfox.app/science/how#paper-cepeda2006)[Cepeda et al. (2008)](https://www.kerfox.app/science/how#paper-cepeda2008)[Kerfoot et al. (2007)](https://www.kerfox.app/science/how#paper-kerfoot2007)[Ye, Su & Cao (2022)](https://www.kerfox.app/science/how#paper-ye2022)[Settles & Meeder (2016)](https://www.kerfox.app/science/how#paper-settles2016)

#### WHAT WE BUILT

Review dates come from FSRS, a scheduler that models your memory of each item as three quantities and works out when it is about to slip. We use the library rather than hand rolling intervals. On top of it sits a clamp taken from that ridgeline result: as an exam approaches, the gap between reviews is pulled toward a fraction of the time left rather than being allowed to run out past it.

#### WHERE YOU MEET IT

Review, and the dates topics come back on. You do not choose them and you are not meant to.

### Answering before re-reading

#### WHAT THE RESEARCH FOUND

Reading something again is the most popular way to revise and one of the least effective. In the study that made the point plainly, students who tested themselves remembered 61 per cent of a passage a week later, while students who reread it four times remembered 40 per cent, despite having read it fourteen times to the testers’ three. The reversal is the interesting part: at five minutes the rereaders were ahead, and they were also the more confident group about what they would remember. A meta-analysis of 159 comparisons puts the durable advantage at a medium effect, larger when the practice asks you to produce the answer rather than pick it.

SOURCES[Roediger & Karpicke (2006)](https://www.kerfox.app/science/how#paper-roediger2006)[Karpicke & Blunt (2011)](https://www.kerfox.app/science/how#paper-karpicke2011)[Rowland (2014)](https://www.kerfox.app/science/how#paper-rowland2014)[Dunlosky et al. (2013)](https://www.kerfox.app/science/how#paper-dunlosky2013)

#### WHAT WE BUILT

There is no re-read button offered as the next step anywhere in the app. When a lesson finishes, what comes next is a question. Most answers are typed rather than chosen, because producing an answer is the part that does the work, and the app draws its own keyboards so that typing a formula on a phone is not the reason someone picks multiple choice.

#### WHERE YOU MEET IT

The end of every lesson, and the whole of Review and Lightning. Memorise goes further and refuses to let a fact go until you have recalled it correctly on separate days.

### Interleaving, and why it has to be the confusable things

#### WHAT THE RESEARCH FOUND

Practising one kind of problem in a block, then the next kind in the next block, feels efficient and tests badly. In a randomised trial across 54 maths classes, students who met the same problems mixed up rather than grouped scored 61 per cent on a surprise test about a month later, against 38 per cent for the blocked group. The meta-analysis explains why and narrows it: the gain comes from having to work out which kind of problem this is, so it appears when the alternatives are easy to confuse and can disappear entirely otherwise. It is also much smaller on average than that trial suggests, and for plain word lists it points the other way.

SOURCES[Rohrer et al. (2020)](https://www.kerfox.app/science/how#paper-rohrer2020)[Brunmair & Richter (2019)](https://www.kerfox.app/science/how#paper-brunmair2019)

#### WHAT WE BUILT

Our interleaver reorders questions only inside clusters of topics a human has explicitly marked as confusable with each other, and leaves everything else in the order it already had. The adjacency is authored, never inferred, because "these two really do get mixed up" is a claim about students that a model has no business making on its own.

#### WHERE YOU MEET IT

Lightning rounds, where a set that would have arrived topic by topic instead puts the two things you mix up next to each other.

### Pitching it where you can just about do it

#### WHAT THE RESEARCH FOUND

A question you are certain to get right teaches almost nothing, and neither does one you cannot possibly do. Two lines of work put numbers on the middle, and they do not agree. A maths practice system used by thousands of Dutch schoolchildren rates the child and the question on a single chess-style scale, updating both after every answer, and aims at about a 75 per cent success rate. Separately, a mathematical result derives roughly 85 per cent as the fastest training accuracy for a broad family of learning algorithms. Those are different kinds of evidence: one is a system that ran in schools, the other is a derivation about machine learners on classification tasks, with no classroom in it and a figure that moves when its assumptions do.

SOURCES[Klinkenberg, Straatemeier & van der Maas (2011)](https://www.kerfox.app/science/how#paper-klinkenberg2011)[Wilson et al. (2019)](https://www.kerfox.app/science/how#paper-wilson2019)

#### WHAT WE BUILT

Your ability is tracked with the same kind of rating, and the serving layer tilts a session toward items near a target success rate: about 0.60 when the job is to measure what you know, 0.85 when the job is to teach, and falling toward 0.70 in the fortnight before a paper, when the job is to find what will break under pressure. It is a weight and never a filter, because a filter on a thin bank produces an empty session, and an empty session is worse than a badly pitched one.

#### WHERE YOU MEET IT

Any session you start: which questions come up, and in what order.

### Difficulties worth keeping

#### WHAT THE RESEARCH FOUND

Some things that make studying feel faster make the learning worse, and some things that make it feel slower and more awkward make it better. That is the argument behind desirable difficulties, and one of its cleaner demonstrations is guessing first: students quizzed on a passage before reading it scored higher afterwards than students simply given extra time to read, and the benefit held even when the count was restricted to the questions they had got wrong on the pretest. Being wrong first was not the cost. It was part of the mechanism.

SOURCES[Bjork & Bjork (2011)](https://www.kerfox.app/science/how#paper-bjork2011)[Richland, Kornell & Kao (2009)](https://www.kerfox.app/science/how#paper-richland2009)

#### WHAT WE BUILT

A solution never arrives as one block. It opens in order: the boxed answer, then the working one step at a time, then a deeper explanation of why the method works, which stays shut unless you open it. And in some lessons a figure asks you to commit to a prediction before the slide that explains it. That guess is never scored, carries no points and cannot move your mastery, because a prediction that costs marks is a test, and the whole thing depends on you being willing to be wrong.

#### WHERE YOU MEET IT

Predictions inside lessons, and the moment after you submit a long question, when the answer appears before the working does.

### Worked examples, faded, and knowing when to stop showing them

#### WHAT THE RESEARCH FOUND

Being shown a full solution and then trying one yourself beats grinding through problems unaided, by about half a standard deviation across 55 studies of maths learning. Two things qualify it. The benefit in the original experiments was specific to problems of the same structure and did not carry to differently built ones. And the effect reverses with expertise: support that clearly helps a beginner becomes redundant for someone who knows the material, and past a point makes their performance worse. The repair is to fade the example out, blanking the last step, then the last two, then handing over the bare problem, which beat alternating examples and problems in a controlled comparison.

SOURCES[Sweller & Cooper (1985)](https://www.kerfox.app/science/how#paper-sweller1985)[Barbieri et al. (2023)](https://www.kerfox.app/science/how#paper-barbieri2023)[Kalyuga et al. (2003)](https://www.kerfox.app/science/how#paper-kalyuga2003)[Atkinson, Renkl & Merrill (2003)](https://www.kerfox.app/science/how#paper-atkinson2003)

#### WHAT WE BUILT

Every long question in the bank was authored with its solution written one step per line, so a faded example is that working with the last steps withheld, and no new authoring was needed for it. There are four rungs, from studying the whole thing to solving it cold, and the fading is backward because that is the direction the study tested. The deepest layer of a solution is opt in for the same reason: an expert should not have to read past it.

#### WHERE YOU MEET IT

The worked example screen, and long questions in a session, where the same problem can arrive as a demonstration or as a blank sheet depending on how far along you are.

### Explaining it back

#### WHAT THE RESEARCH FOUND

Producing an answer yourself is remembered better than reading the same answer, across 445 comparisons, though most of that work was done on single words in a laboratory. Explaining the steps of a worked example to yourself, rather than reading it through, is what separated the students who learned from a physics example from the ones who did not, and prompting people to explain reliably helps across subjects. Simply expecting to have to teach something later makes recall fuller and better organised than expecting to be tested on it. One result cuts the other way and belongs here: the maths meta-analysis above found that bolting self-explanation prompts onto worked examples made them slightly worse, so this is not a technique that improves everything it is added to.

SOURCES[Slamecka & Graf (1978)](https://www.kerfox.app/science/how#paper-slamecka1978)[Bertsch et al. (2007)](https://www.kerfox.app/science/how#paper-bertsch2007)[Chi et al. (1989)](https://www.kerfox.app/science/how#paper-chi1989)[Bisra et al. (2018)](https://www.kerfox.app/science/how#paper-bisra2018)[Nestojko et al. (2014)](https://www.kerfox.app/science/how#paper-nestojko2014)

#### WHAT WE BUILT

Teach It asks you to explain a topic in your own words and marks what you wrote by checking whether each idea the explanation needs is actually present, and separately whether you have said two things that cannot both be true. Length never enters the score. A contradiction is flagged harder than an omission, because a board that finds a right point and a wrong one alongside it marks the pair wrong.

#### WHERE YOU MEET IT

Teach It, reached from a topic you have already worked through.

### Feedback that says which mark and why

#### WHAT THE RESEARCH FOUND

Feedback is among the strongest influences on achievement and also one of the most variable: a synthesis of twelve meta-analyses puts its average effect at about twice a year of ordinary schooling, while praise aimed at the person barely registers. It can also do harm. Pooling 607 effects, more than a third of feedback interventions made performance worse, and they did so as attention moved away from the task and toward the self. The reviews agree on the shape that works: about the task, specific, and explaining the what and the why rather than confirming right or wrong. One experiment sharpens it usefully. Being told why an answer was right, rather than just what it was, made no difference at all on the same questions two days later, and made the whole difference on new questions needing an inference.

SOURCES[Hattie & Timperley (2007)](https://www.kerfox.app/science/how#paper-hattie2007)[Shute (2008)](https://www.kerfox.app/science/how#paper-shute2008)[Kluger & DeNisi (1996)](https://www.kerfox.app/science/how#paper-kluger1996)[Butler, Godbole & Marsh (2013)](https://www.kerfox.app/science/how#paper-butler2013)

#### WHAT WE BUILT

Kerfox does not mark answers right or wrong. It marks them out of marks, and the border of the answer box is cut into one arc per mark, each landing earned or dropped, so three out of four is readable without a number. Where a wrong answer is a recognisable slip, the app names it: a missing one half, degrees where radians belong, a radius read as a diameter. Naming it never changes the mark, which is what makes it safe for the app to have a guess at all. Marks carried forward from your own earlier error are credited the way an examiner would credit them.

#### WHERE YOU MEET IT

The moment you submit any part of a long question.

### Figures that are simulated rather than drawn

#### WHAT THE RESEARCH FOUND

Adding a diagram to a passage helps readers build a working mental model and draw inferences from it, not just remember more sentences. The detail that shaped our house style is which diagram won: simplified drawings beat anatomically detailed ones, both for learning the facts and for tying them together. Words and pictures also have to arrive together rather than one after the other. On whether a moving picture beats a still one, the honest answer is that it depends and the margin is modest. A 2007 meta-analysis put animation ahead by a small to medium amount, with the large gains confined to realistic video and to physical procedures, and later and larger reviews put the average lower still.

SOURCES[Butcher (2006)](https://www.kerfox.app/science/how#paper-butcher2006)[Mayer & Anderson (1992)](https://www.kerfox.app/science/how#paper-mayer1992)[Höffler & Leutner (2007)](https://www.kerfox.app/science/how#paper-hoffler2007)

#### WHAT WE BUILT

Our figures are stripped to the claim they are making, and anything that moves is computed from a function of time or integrated as a real simulation rather than played back from a list of poses. Gas molecules fly in straight lines and exchange velocity along the line of centres, so momentum and energy are exactly conserved and no molecule turns a corner where nothing is there to turn it. The drawing has to be true: if the caption says the steps are the same length, they are generated the same length. Nothing glows, and colour carries meaning rather than decoration.

#### WHERE YOU MEET IT

Every lesson slide that carries a figure, and the play control that only appears on the ones that genuinely loop.

## Things we built and then took out.

A page listing only the decisions that survived is a page that has quietly hidden its method. These are ideas we implemented, or were about to, and then removed. Each one was killed by something specific, and the something is the useful part.

They are not all about learning. Two of them are about a keyboard and one is about how a diagram is drawn, and they are here because the same rule killed all of them: a thing that seems obviously good is worth measuring, and about a third of the time the measurement says stop.

What the planner prescribes

### Flashcard decks as the day's default

#### WHAT WE HAD

A deck system, and both planners prescribing a review of it. It is the obvious build: retrieval practice is the best-evidenced thing in the whole field, and a deck is the cheapest possible delivery of it.

#### WHY IT WENT

The evidence is for retrieval on a schedule somebody else keeps, and a deck hands the schedule to the learner. Given the option to set aside cards they judged they knew, learners did WORSE, because the judgement a deck rests on is the one people are least reliable at. So the retrieval stays and the deck stops being what the day prescribes; the same practice runs to a criterion the app holds instead.

SOURCES[Kornell & Bjork (2008)](https://www.kerfox.app/science/how#paper-kornell2008)[Roediger & Karpicke (2006)](https://www.kerfox.app/science/how#paper-roediger2006)

ON ITS WAY OUT

The pairing engine

### Serving only questions at the right difficulty

#### WHAT WE HAD

A filter on the pool: work out the band a student should be in and serve nothing outside it. Every word of the difficulty research points at it.

#### WHY IT WENT

It empties a session. Our bank is young, and a filter over a thin pool returns three questions on a topic that needed twelve - which is worse than a badly pitched session by a distance no pedagogy closes. It became a weighting instead, which gives the same answer when the pool is rich and degrades to today's behaviour when it is not.

OUT OF THE APP

The marker

### Marking a formula by comparing it to the answer

#### WHAT WE HAD

The obvious implementation, and the one nearly every app ships: normalise the spacing, compare the strings, mark it.

#### WHY IT WENT

It marks NOTATION rather than physics. A student who writes root 2gh where we wrote sqrt(2*g*h) has the right answer and gets nothing, and every rearrangement they are entitled to make is a new way to be wrongly marked wrong. Both expressions are now evaluated at a dozen scattered points and marked equal if they agree everywhere, so the marking is of the function rather than of the typing.

OUT OF THE APP

The marker

### Derivations the student marks themselves

#### WHAT WE HAD

Show the model derivation, ask the student to compare, let them award their own marks. It is what a textbook does and it is honest about the difficulty of the alternative.

#### WHY IT WENT

Self-marking rewards whoever is hardest on themselves, which is the wrong student. A derivation is now marked line by line: each line is tested for whether it is TRUE given the question's physics, and then for which relations it actually used, which is what separates real work from writing the final answer down. The bank is still migrating, so some derivations still self-mark today.

ON ITS WAY OUT

The keyboards

### A keyboard of our own for chemistry

#### WHAT WE HAD

A grid of forty element keys. Typing NaCl was two taps instead of five, and for an equation full of formulae the saving compounds.

#### WHY IT WENT

It was a fifth keyboard. Faster for the chemistry, and a whole surface a student had to learn separately from the one they already knew - and the sums only worked for the students who had already found it. Chemistry types on the same QWERTY as everything else now, with a strip above the keys carrying the arrows, this question's own elements and the state symbols.

OUT OF THE APP

The keyboards

### A buzz on every key

#### WHAT WE HAD

Haptic feedback on every keystroke. It feels right for about ten seconds, and it is what a physical keyboard would do.

#### WHY IT WENT

A sentence is a hundred keystrokes, so it is a phone humming continuously in a hand for a minute, and a signal that fires on everything says nothing. It is spent now on the two keys that do something other than put a character on the screen: shift, whose effect only arrives a keystroke later, and backspace, which takes something away and is the one key people press without looking.

OUT OF THE APP

The figures

### Glow on the figures

#### WHAT WE HAD

A soft halo under every beam, every photon and every hot body. Individually each one looked better than the flat version.

#### WHY IT WENT

Together they made five hundred diagrams look like five hundred screensavers. A soft copy under a line does not add light, it softens the edge of a line drawn crisp on purpose, and once everything wears one the glow has stopped distinguishing anything. Depth comes from a gradient inside a solid body now, and from nothing else.

OUT OF THE APP

The figures

### Keyframed loops in the animations

#### WHAT WE HAD

Five or six poses per loop, interpolated. It is how almost every web animation is built and it is far less code.

#### WHY IT WENT

A reader can see both failures. It is STEPPED, because five keyframes over three seconds is five poses and four interpolations, each eased separately, so a wave crawls between postures instead of oscillating. And it has a SEAM: the list ends where it began, so a travelling thing rewinds and the eye catches the rewind every single cycle. Anything that repeats is sampled from a function of time now, and a loop built from a sine has no place where anything happens.

OUT OF THE APP
