AceQuiz LogoAceQuiz

The learning science behind AceQuiz

AceQuiz is built on four findings from published education research. Here they are, in the units the papers actually use, with links so you can check every one.

● Evidence page · last reviewed July 30, 2026

Study tools make a lot of claims about memory. Most of them are unsourced. Our rule is simple: if a number appears anywhere on this site, it appears on this page first, with the paper it came from. Where the research is strong we say so. Where it does not stretch as far as a marketing team would like, we say that too.

4 findings · 2 meta-analyseseffect sizes, not invented multipliersevery claim mapped to a shipped feature

Four findings we build on

Two of these are meta-analyses, which pool many experiments rather than resting on a single result. That is the strongest form this evidence comes in.

Finding 01Review of Educational Research, 87(3)

Adesope, Trevisan & Sundararajan (2017)

Pooling decades of experiments, students who practiced by testing themselves outperformed students who reread the same material.

Practice testing vs. rereadingg = 0.51
Practice testing vs. no review at allg = 0.93
Multiple choice practice formatg = 0.70
Short answer practice formatg = 0.48

Measured in standard deviations (Hedges’s g), not percentage points. A g of 0.51 means the average tested student scored about half a standard deviation above the average rereader.

In AceQuiz

Every tool ends in practice you take, not notes you look at. Quizzes, flashcards and mock exams are generated from your own uploads.

Read the source paper ↗
Finding 02Psychological Bulletin, 140(6)

Rowland (2014)

A meta-analysis of 61 experiments found that the benefit of testing nearly doubles when the test tells you why you were wrong.

Testing with feedbackg = 0.73
Testing without feedbackg = 0.39

Same scale as above. The gap between the two bars is the entire argument for explaining answers instead of just scoring them.

In AceQuiz

Every question carries a worked explanation, and typed answers are graded with written feedback rather than marked right or wrong.

Read the source paper ↗
Finding 03Psychological Science in the Public Interest, 14(1)

Dunlosky, Rawson, Marsh, Nathan & Willingham (2013)

Ten common study techniques were rated for real-world utility. Only two earned a high rating: practice testing and distributed practice.

This paper rates techniques rather than reporting a single effect size. The ratings are reproduced in full below.

In AceQuiz

The two high-utility techniques are the two the product is organized around. Summaries and notes exist as inputs to practice, never as the finish line.

Read the source paper ↗
Finding 04Psychological Science, 19(11)

Cepeda, Vul, Rohrer, Wixted & Pashler (2008)

With more than 1,350 participants tested up to a year later, the best gap between study sessions was not fixed. It scaled with how far away the test was.

Optimal gap fell from roughly 20 to 40 percent of a one-week retention interval down to roughly 5 to 10 percent of a one-year interval.

In AceQuiz

Your exam date is a scheduling input, not decoration. Smart Review and the study plan compress or stretch around the day you actually sit the paper.

Read the source paper ↗

All ten techniques, rated. Two of them made the top tier

Dunlosky et al. (2013) graded the study methods students actually use. The habits most people default to sit in the bottom tier.

High utility2

Works across ages, subjects and test formats.

  • Practice testing
  • Distributed practice
Moderate utility3

Real benefits, but narrower conditions.

  • Elaborative interrogation
  • Self-explanation
  • Interleaved practice
Low utility5

The techniques students reach for most often.

  • Summarization
  • Highlighting and underlining
  • Keyword mnemonic
  • Imagery for text
  • Rereading

Source: Dunlosky, Rawson, Marsh, Nathan & Willingham (2013), Psychological Science in the Public Interest 14(1). Ratings reproduced in full, including the tiers that are inconvenient for a company that also ships a summarizer.

How to read an effect size

This is the step most study tools skip, because skipping it is how a 0.5 becomes a headline.

What g measures

Hedges’s g is a difference measured in standard deviations, not in marks. A g of 0.51 says the average student who practiced by testing scored about half a standard deviation above the average student who reread. It does not say plus 51 percent, and anyone reporting it that way has made a mistake or a choice.

What that looks like on a real paper

It depends entirely on how spread out your class is. In a course where the mean is 60 and scores are typically 12 marks apart, 0.51 standard deviations is close to 6 marks, which is roughly a 10 percent improvement on that student’s own result. In a tightly bunched class it is worth less. In a wide one it is worth more.

What we do not claim
  • Not our results. These studies test methods, not this product. No published trial has measured AceQuiz.
  • No guaranteed grade. Effect sizes describe averages across groups. They do not predict one person’s exam.
  • No invented multipliers. You will not find a times-faster figure here, because none of these papers report one.
  • More practice is not always better. Adesope found a single well-timed practice test beat several stacked together.

The scheduler, and the papers under it

Spacing only helps if something decides the intervals. Ours is a published memory model, not a rule of thumb.

Smart Review

Smart Review estimates how likely you are to still recall each card and schedules the next sighting just before that probability falls away. It runs on FSRS (Free Spaced Repetition Scheduler), a difficulty / stability / retrievability memory model. Because it models your ratings rather than counting days, two students working the same deck get different queues.

Your exam date then bends the whole schedule. That is the Cepeda finding applied: the right gap is a fraction of the time you have left, so a paper in nine days and a paper in nine weeks should not produce the same plan.

Peer-reviewed source
  • A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling

    ACM SIGKDD 2022

  • Optimizing Spaced Repetition Schedule by Capturing the Dynamics of Memory

    IEEE Transactions on Knowledge and Data Engineering

The open-source model ↗

What that looks like in the product

One loop, four steps, three of them lifted straight from the papers above.

  1. Step 01

    Generate from your own material

    Slides, PDFs, recordings and YouTube links become questions. This step is logistics, not learning. It exists so the next three can happen at all.

  2. Step 02Adesope et al., 2017

    Retrieve, do not reread

    You answer from memory before you look anything up. This is the move that carries the largest measured benefit of the four.

  3. Step 03Rowland, 2014

    Find out why you were wrong

    Each answer comes back with a worked explanation, and typed answers get written feedback. Skipping this throws away about half the effect.

  4. Step 04Cepeda et al., 2008

    Come back at the right distance

    Weak concepts return on a schedule set by your ratings and your exam date, rather than whenever you happen to open the app.

Run the loop on your own material

Upload a lecture deck and take a practice round in about two minutes. The free tier does not expire and does not ask for a card.

Questions about the evidence

Does practice testing actually raise exam scores?
Yes, and it is one of the better-replicated findings in education research. The largest meta-analysis of practice testing (Adesope et al., 2017, Review of Educational Research) pooled experiments spanning four decades and found students who tested themselves outperformed students who reread the same material by 0.51 standard deviations. Against no review at all the gap was 0.93. Effect sizes are not percentages, so what that means for your grade depends on how spread out your class scores are.
Why does AceQuiz explain every answer instead of just scoring me?
Because feedback is where most of the benefit lives. Rowland's 2014 meta-analysis in Psychological Bulletin found that testing with feedback produced an effect of 0.73 standard deviations, against 0.39 for testing without it. Scoring you and moving on throws away roughly half the value of the practice, so every question carries a worked explanation and typed answers are graded with written feedback.
Is rereading and highlighting really that bad?
They are not harmful, they are just weak per hour spent. Dunlosky et al. (2013) reviewed ten common study techniques and rated only two as high utility: practice testing and distributed practice. Rereading, highlighting, summarization, keyword mnemonics and imagery all landed in the low utility tier. That is why AceQuiz treats notes and summaries as inputs to practice rather than as the finished study session.
How does AceQuiz decide when to show me a card again?
Smart Review schedules each card by modeling how likely you are to still recall it, using the FSRS memory model published at ACM SIGKDD and IEEE TKDE. The interval is not fixed, it adapts to your rating history and to how far away your exam is. Cepeda et al. (2008), testing more than 1,350 people, found the best gap between sessions scales with the retention interval, from roughly 20 to 40 percent of a one-week horizon down to 5 to 10 percent of a one-year horizon, which is why your exam date changes your schedule.
Do you claim AceQuiz makes people learn a specific percentage faster?
No. We publish the effect sizes the studies report, in the units the studies use, and we do not invent multipliers. The research supports the methods AceQuiz is built on. It does not measure AceQuiz itself, and we say so on this page rather than dressing published findings up as our own results.