Kerfox

THE RESEARCH BEHIND KERFOX

Why we built it this way.

We have no evidence that Kerfox works. Nobody has run a study on it, no class has been taught with it and measured against a class that was not, and until somebody does, any number we put on this page would be one we made up.

What we do have is this. Every design decision in the app was taken from a published finding, and we can show you which one. Below, each method is named, the research behind it is cited and linked, and we say exactly where in the app you run into it, so you can go and look rather than take our word for it.

That is a claim about our methods rather than about your marks. It is the only honest claim available to a study app that nobody has studied yet, and we would rather make it well than make a bigger one badly.

  • 33papers read, linked and summarised
  • 9methods, each traced back to them
  • 42checks that fail the build rather than warn
  • 0studies of our own

The last one is the number that decides how the rest of this page is written.

WHAT THIS PAGE IS NOT

It is not an efficacy page. The findings below were established by other people, in their own studies, on their own materials. That a technique works in a published experiment is not evidence that our implementation of it works, and we are careful throughout to keep those two things apart.

It is also not a sales brochure with footnotes. Where a finding is contested, or where its scope is narrower than the headline suggests, we say so in the same paragraph rather than quietly rounding it up. Three of the methods below carry a sentence that cuts against them.

What we do not know, what a study of Kerfox would have to look like, and every paper this app was built on.

What we do not know yet.

The findings above are real and they are not ours. Building on them is a reasonable bet, and a bet is exactly what it is. Here is what we genuinely cannot tell you.

  1. Whether any of it helps.

    The honest headline. Every finding above was established on other people’s materials with other people’s learners, and none of it tells you what happens when a student opens our app in October and sits an exam in June.

  2. How much of a laboratory result survives the walk to a phone.

    A spacing effect measured on word lists under supervision is not the same object as a spacing effect on A-Level chemistry, on a bus, at eleven at night, competing with everything else on the screen.

  3. What the right difficulty actually is.

    The theory that motivates our target is a derivation about machine learners, and the one classroom system that published its number chose a different one. We picked a band, we can say exactly where it came from, and we cannot tell you it is right.

  4. Whether making a student teach a topic back beats simply asking them more questions about it.

    The research is clear that generating beats reading. It says much less about generating versus being tested, which is the comparison our app actually puts in front of someone choosing what to do next.

  5. What our marking gets wrong.

    At GCSE and A-Level we mark against particular words because the boards do, and we are sometimes deliberately kinder than they would be. We say so where we know it. We do not have the full list.

  6. Whether the whole approach suits everyone.

    Most of this research was done on students who were already doing reasonably well in fairly conventional settings. Whether it holds for a student who is behind, anxious, or working in a second language is a question the literature answers much less confidently than its summaries suggest.

Our own studies.

This is where a study of Kerfox would go. There is not one, so the section is empty, and it is going to stay visibly empty rather than quietly disappear until there is something true to put in it.

0

studies of Kerfox, by us or by anybody

WHAT WILL BE HERE

  1. The question we set out to answer, and the protocol, written down BEFORE we look at any data.
  2. Who took part, how many, for how long, and every way they are not a random sample of students.
  3. The numbers themselves, in a form somebody else can re-analyse.
  4. The result whichever way it goes, including the version where our own feature made no difference.

WHAT WILL NOT BE

  1. An in-house A/B test presented as evidence that Kerfox works. It would measure which of two Kerfoxes is better, which is a different question and a much easier one.
  2. Engagement numbers standing in for learning. Time in the app is the metric it is easiest to move and the one it means least to move.
  3. Testimonials. A quote from somebody who liked it is not a finding, and printing it next to real citations would borrow their credibility for it.

AN INVITATION

If you teach, run a department, or do research in education and you would like to find out whether any of this works, we would like that very much. We can give you the app free, the data a study would need, and whatever build you want to test against. We will publish what you find whichever way it goes, and we will say so here if it goes against us.

Talk to us about a study

The reading list.

Every paper named above, in full, each one with what it actually showed and which of our methods leans on it. Where a link goes to a publisher who wants money for the article, the authors have very often put a copy on their own university page, and it is worth looking.

All 33 papers.

Every paper this page leans on, with what it showed and which part of Kerfox it is behind.
What it showedBehind
Murre & DrosReplication and Analysis of Ebbinghaus' Forgetting CurvePLOS ONE, 10(7), e0120644doi.org/10.1371/journal.pone.01206442015A faithful modern repeat of Ebbinghaus’s original self-experiment reproduced his forgetting curve: recall of nonsense syllables collapses within the first hours and then flattens. Like the original it has one subject, and the authors note the curve is not perfectly smooth, ticking up again after 24 hours in a way they connect to sleep.The schedulerSpacing
Cepeda et al.Distributed practice in verbal recall tasks: A review and quantitative synthesisPsychological Bulletin, 132(3), 354-380doi.org/10.1037/0033-2909.132.3.3542006Across a century of experiments, spaced study beat massed study on the final test, 47 per cent against 37 per cent, and only 12 of 271 comparisons failed to show the benefit. The materials are overwhelmingly word lists and paired associates rather than school subjects.Spacing
Cepeda et al.Spacing effects in learning: A temporal ridgeline of optimal retentionPsychological Science, 19(11), 1095-1102doi.org/10.1111/j.1467-9280.2008.02209.x2008With 1,354 people, gaps up to three and a half months and tests up to a year later, there is a best gap between study sessions for any given deadline, and it is a shrinking fraction of that deadline: about a fifth of it at short delays, nearer a twentieth at a year.The schedulerSpacing
Kerfoot et al.Spaced education improves the retention of clinical knowledge by medical students: a randomised controlled trialMedical Education, 41(1), 23-31doi.org/10.1111/j.1365-2929.2006.02644.x2007Medical students were emailed a weekly question on half of their topics and nothing on the other half, making each student their own control, and at the end of the year they scored higher on exactly the topics they had been drip fed. The gap was widest for material first studied nine to eleven months earlier.Spacing
Ye, Su & CaoA Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition SchedulingProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4381-4390doi.org/10.1145/3534678.35390812022From 220 million real review logs, a memory model was fitted and review scheduling recast as a shortest-path problem, cutting the work needed to memorise a word by about 13 per cent against the best previous schedulers. The FSRS scheduler we use descends from this model; FSRS itself has no peer-reviewed paper of its own, and we would rather say that than imply one exists.The schedulerSpacing
Settles & MeederA Trainable Spaced Repetition Model for Language LearningProceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1848-1858aclanthology.org/P16-1174/2016Duolingo trained a model on 13 million real review traces to predict how long a word stays in memory, cutting recall-prediction error by more than 45 per cent against standard heuristics and raising daily engagement by 12 per cent in a live trial. It is the clearest published evidence that a learned scheduler beats a fixed one at scale.Spacing
Roediger & KarpickeTest-Enhanced Learning: Taking Memory Tests Improves Long-Term RetentionPsychological Science, 17(3), 249-255doi.org/10.1111/j.1467-9280.2006.01693.x2006Students who tested themselves recalled 61 per cent of a passage a week later against 40 per cent for students who reread it four times, and the ordering was the other way round at five minutes. The rereaders were also the more confident group about what they would remember.RetrievalFlashcard decks as the day's default
Karpicke & BluntRetrieval Practice Produces More Learning than Elaborative Studying with Concept MappingScience, 331(6018), 772-775doi.org/10.1126/science.11993272011With study time held equal, students who practised recalling a science text scored about half again as much a week later as students who built concept maps from it, on inference questions as well as recall, and they had predicted the opposite. A formal Comment was published on this paper in the same journal; it did not overturn the result, and it is worth knowing the paper is argued over.Retrieval
RowlandThe Effect of Testing Versus Restudy on Retention: A Meta-Analytic Review of the Testing EffectPsychological Bulletin, 140(6), 1432-1463doi.org/10.1037/a00375592014Pooling 159 comparisons, testing yourself beat rereading by a medium margin, g = 0.50, with the advantage larger when the practice test asked for recall rather than recognition. The spread between studies is very wide.Retrieval
Dunlosky et al.Improving Students’ Learning With Effective Learning Techniques: Promising Directions From Cognitive and Educational PsychologyPsychological Science in the Public Interest, 14(1), 4-58doi.org/10.1177/15291006124532662013A review rating ten study techniques gave its highest recommendation to only two, testing yourself and spreading study out, while highlighting and rereading came out among the weakest. Interleaving and self-explanation were rated moderate rather than high, which is often misquoted upward.Retrieval
Rohrer et al.A randomized controlled trial of interleaved mathematics practiceJournal of Educational Psychology, 112(1), 40-52doi.org/10.1037/edu00003672020In 54 seventh-grade maths classes, students who practised the same problems mixed up scored 61 per cent on an unannounced test about a month later against 38 per cent for students who practised them grouped by topic. The authors note the honest framing is a low against a high dose of interleaving, since every class also got some of each.Interleaving
Brunmair & RichterSimilarity matters: A meta-analysis of interleaved learning and its moderatorsPsychological Bulletin, 145(11), 1029-1052doi.org/10.1037/bul00002092019Across 238 effect sizes interleaving helped moderately overall, most when the categories looked alike from outside and varied within, so that telling them apart is the hard part. It splits sharply by material: paintings did well, maths came out at less than half the size of the trial above, and word lists actually favoured blocking.Interleaving
Klinkenberg, Straatemeier & van der MaasComputer adaptive practice of Maths ability using a new item response model for on the fly ability and difficulty estimationComputers & Education, 57(2), 1813-1824doi.org/10.1016/j.compedu.2011.02.0032011The Maths Garden system rates the child and the question on one chess-style scale, updating both after every answer using speed as well as accuracy, and serves questions at a mean success probability of 0.75. Over ten months 3,648 children answered more than three and a half million problems.Difficulty
Wilson et al.The Eighty Five Percent Rule for optimal learningNature Communications, 10, 4646doi.org/10.1038/s41467-019-12552-42019For a broad family of algorithms that improve by small steps after each mistake, learning is fastest when training is tuned so the learner is right about 85 per cent of the time. This is a derivation demonstrated on neural networks and a model of animal perceptual learning, with no human classroom in it, and the figure moves to 82 or 75 per cent under different noise assumptions.The pairingDifficulty
Bjork & BjorkMaking Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance LearningIn Gernsbacher et al. (Eds.), Psychology and the Real World, chapter 5, 56-64. Worth Publishersbjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf2011Conditions that make performance improve quickly during study often fail to support later retention and transfer, while conditions that feel harder and slower at the time often improve both. It is an argued essay drawing together decades of experiments rather than a single study.The pairingDesirable difficulty
Richland, Kornell & KaoThe pretesting effect: Do unsuccessful retrieval attempts enhance learning?Journal of Experimental Psychology: Applied, 15(3), 243-257doi.org/10.1037/a00164962009Across five experiments, students quizzed on a passage before reading it scored higher afterwards than students simply given more time to study, and the effect held when the count was restricted to questions they had got wrong on the pretest. The materials are real educational prose, which is why we prefer it to the tidier word-pair versions.Desirable difficulty
Sweller & CooperThe Use of Worked Examples as a Substitute for Problem Solving in Learning AlgebraCognition and Instruction, 2(1), 59-89doi.org/10.1207/s1532690xci0201_31985Students who studied worked algebra examples spent much less time learning and then solved later problems faster and with fewer errors than students who worked them out unaided. Both benefits were specific to problems built the same way, and did not carry to structurally different ones.Worked examples
Barbieri et al.A Meta-analysis of the Worked Examples Effect on Mathematics PerformanceEducational Psychology Review, 35, article 11doi.org/10.1007/s10648-023-09745-12023Across 55 studies from primary school to adulthood, studying worked solutions came out about half a standard deviation ahead of practice alone, g = 0.48, holding at 0.44 after correcting for publication bias, which the authors did detect. It also found that adding self-explanation prompts made worked examples slightly worse, which cuts against a technique we use elsewhere.Worked examples
Kalyuga et al.The Expertise Reversal EffectEducational Psychologist, 38(1), 23-31doi.org/10.1207/s15326985ep3801_42003Teaching support that clearly helps a beginner loses its power as the learner gains knowledge, and past a point makes performance worse than no support at all, so guidance has to be withdrawn as expertise grows. It is a review of existing experiments rather than a new one.Worked examples
Atkinson, Renkl & MerrillTransitioning From Studying Examples to Solving Problems: Effects of Self-Explanation Prompts and Fading Worked-Out StepsJournal of Educational Psychology, 95(4), 774-783doi.org/10.1037/0022-0663.95.4.7742003Learners moved from a complete worked solution to the same problem with the last step blanked, then the last two, then the bare problem, and solved more new problems afterwards than learners given alternating examples and problems, at no extra time. The comparison against non-faded practice rests on one experiment with 78 university students working on probability.Worked examples
Slamecka & GrafThe Generation Effect: Delineation of a PhenomenonJournal of Experimental Psychology: Human Learning and Memory, 4(6), 592-604doi.org/10.1037/0278-7393.4.6.5921978Words people produced themselves from a rule were remembered better than the identical words simply read, across recognition, free recall and cued recall. This is single-word laboratory memory, not comprehension of a subject.Generation
Bertsch et al.The Generation Effect: A Meta-Analytic ReviewMemory & Cognition, 35(2), 201-210doi.org/10.3758/BF031934412007Summarising 445 effect sizes across 86 studies, generating information rather than reading it gave a memory benefit of about d = 0.40, almost half a standard deviation. The size varied substantially with the conditions, so 0.40 is an average and not a promise.Generation
Chi et al.Self-Explanations: How Students Study and Use Examples in Learning to Solve ProblemsCognitive Science, 13(2), 145-182doi.org/10.1207/s15516709cog1302_11989The students who learned most from worked physics examples were the ones who talked themselves through why each step worked; the weaker students explained little, overestimated their own understanding, and leaned on copying the example. It is a small, intensive study of university students on mechanics.Generation
Bisra et al.Inducing Self-Explanation: a Meta-AnalysisEducational Psychology Review, 30(3), 703-725doi.org/10.1007/s10648-018-9434-x2018Across 69 effect sizes from 64 studies, prompting learners to explain material to themselves improved learning by about g = 0.55, holding across a wide range of subjects and tasks. The prompts were written by instructors, which the authors flag as a limitation.Generation
Nestojko et al.Expecting to teach enhances learning and organization of knowledge in free recall of text passagesMemory & Cognition, 42(7), 1038-1048doi.org/10.3758/s13421-014-0416-z2014People told they would later teach a passage recalled it more fully and in a better organised way than people told they would be tested, even though nobody taught anyone and everyone was tested. The advantage on direct questions showed up specifically on the passage’s main points.Generation
Hattie & TimperleyThe Power of FeedbackReview of Educational Research, 77(1), 81-112doi.org/10.3102/0034654302984872007Drawing twelve earlier meta-analyses together, feedback averaged an effect of 0.79, roughly twice that of an ordinary year of schooling, but the average hides a wide range: information about the task and how to do it better works, while praise and rewards aimed at the person sit near 0.14.The mark schemeFeedback
ShuteFocus on Formative FeedbackReview of Educational Research, 78(1), 153-189doi.org/10.3102/00346543073137952008A review of decades of feedback research concluding that feedback should be about the task rather than the learner, specific, and should explain the what, how and why rather than confirming right or wrong, with praise used sparingly if at all. It states plainly that there is no single best kind.The mark schemeFeedback
Kluger & DeNisiThe effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theoryPsychological Bulletin, 119(2), 254-284doi.org/10.1037/0033-2909.119.2.2541996Pooling 607 effects, feedback helped on average but more than a third of interventions made performance worse, and effectiveness fell as attention moved away from the task and toward the self. The corpus is largely workplace and laboratory performance rather than classroom learning.Feedback
Butler, Godbole & MarshExplanation feedback is better than correct answer feedback for promoting transfer of learningJournal of Educational Psychology, 105(2), 290-298doi.org/10.1037/a00310262013Students given an explanation of why an answer was right did no better than students given the answer alone when the same questions came back two days later, and clearly better on new questions requiring an inference. The whole benefit was in the transfer, which is a narrower and more useful claim than "explanations help".The mark schemeFeedback
ButcherLearning from text with diagrams: Promoting mental model development and inference generationJournal of Educational Psychology, 98(1), 182-197doi.org/10.1037/0022-0663.98.1.1822006Students learning about the heart built better mental models and drew more correct inferences when diagrams accompanied the text, and simplified diagrams beat anatomically detailed ones both for the facts and for tying them together. One topic, one group of university students.Figures
Mayer & AndersonThe instructive animation: Helping students build connections between words and pictures in multimedia learningJournal of Educational Psychology, 84(4), 444-452doi.org/10.1037/0022-0663.84.4.4441992Students who saw an animation while hearing its narration solved new problems about it better than students who got the same animation and narration one after the other. The advantage was on problem solving only; plain recall did not separate the groups.Figures
Höffler & LeutnerInstructional animation versus static pictures: A meta-analysisLearning and Instruction, 17(6), 722-738doi.org/10.1016/j.learninstruc.2007.09.0132007Across 26 studies animation beat static pictures by a small to medium margin, d = 0.37, with much larger gains confined to realistic video and to learning a physical procedure. Later and larger meta-analyses put the average lower, so we treat this as the optimistic end of a contested literature rather than as settled.Figures
Kornell & BjorkOptimising self-regulated study: The benefits - and costs - of dropping flashcardsMemory, 16(2), 125-136doi.org/10.1080/096582107017638992008Letting learners set aside the cards they judged they had already learned produced small but consistent DECREASES in later recall against keeping the whole set in rotation. The authors put the failure in the judgement itself: dropping has a compelling logic, and it is only ever as good as a learner’s sense of what they actually know, which is the part that is unreliable.Flashcard decks as the day's default