How often should you
review vocabulary?
Sooner than feels necessary, then further apart each time. The gap that works scales with how long you want to keep the word: to still know it a week later, review after about a day; to still know it a year later, wait about three weeks (Cepeda et al., 2008). Across 254 studies, spreading the same practice out produced 47% recall against 37% for cramming (Cepeda et al., 2006). The schedule is simple to state and almost impossible to run by hand — which is the real problem, and the one worth automating.
Every figure below is sourced at the end of the page
What the best gap actually is
Most advice about review frequency is a number somebody liked the sound of. There is a better source: Cepeda, Vul, Rohrer, Wixted and Pashler (2008) taught 1,354 people a set of obscure facts, then varied two things independently — the gap before the single review, from the same day out to 105 days, and the delay before the final test, from 7 days out to 350. That design answers the question directly, because it lets the optimal gap fall out of the data instead of being assumed.
The result is the single most useful finding in the literature for anyone building a review habit: the longer you want to remember something, the longer you should wait before reviewing it — but the gap grows more slowly than the target does, so it becomes a shrinking share of it.
| You want to still know it in | Best gap to the next review | Share of the target | What the study found |
|---|---|---|---|
| 1 week | about 1 day | ~14% | Reviewing the same day scored worst of every gap tested |
| 5 weeks | about 11 days | ~31% | Gaps of a day or less lost roughly a third of the achievable score |
| 10 weeks | about 3 weeks | ~30% | The peak is broad — anything from about a week to two months held up well |
| 1 year | about 3 weeks | ~6% | Longer gaps still beat short ones, but the optimum stops climbing |
| The rule of thumb | longer target → longer gap | ~5–30% | Being roughly right beats being precise; being spaced at all beats both |
Read these as a ridgeline, not a recipe. The authors' own framing is that retention as a function of gap rises to a broad peak and then declines slowly — so a review that is a few days off the optimum costs very little, while no review at all costs everything. The practical translation is an expanding ladder: first review within a day, second within a week, then intervals that roughly double each time you get the word right.
Why “review it when you feel shaky” fails
The obvious schedule is to review a word when you notice you are losing it. It does not work, and the reason is not laziness — it is that the feeling of knowing and the fact of knowing come apart, in a specific and well-documented direction.
- Familiarity is mistaken for recall. A word you have just seen feels known because it is still in working memory. Roediger and Karpicke (2006) found students judged rereading to be the more effective strategy while performing worse on the delayed test than students who had been tested instead. The strategy that feels productive is reliably the weaker one.
- Forgetting is fastest on day one. Murre and Dros (2015) re-ran Ebbinghaus' original self-experiment and reproduced the shape: the steepest losses happen within the first hours and the first day, then the curve flattens. By the time a word feels shaky, most of the decay has already happened — you are not reviewing, you are relearning.
- Easy reviews barely count. A review that costs no effort adds little. Bjork and Bjork (2011) named this class of effects desirable difficulties: conditions that slow you down in the moment and improve retention afterwards. A schedule built on comfort systematically removes the difficulty that was doing the work.
How many reviews a single word needs before any of this stops mattering: how many times you need to see a word.
The common schedules, compared
Every review system is an answer to one question: what decides when a word comes back? The systems below are ordered roughly by how much of that decision they take off your hands. The column that predicts whether a schedule survives contact with a real month is not the quality of its algorithm — it is what has to happen for the review to occur at all.
| Schedule | What sets the interval | What it asks of you daily | Where it breaks |
|---|---|---|---|
| Cramming before a test | Nothing — one sitting | Nothing, then hours | Works for a Friday test; the 254-study synthesis puts it ten points behind spacing on delayed recall |
| Review everything, daily | The size of your list | Grows without limit | Collapses under its own weight at a few hundred words — and reviews known words at their least useful moment |
| Calendar reminders | You, in advance | A decision plus a session | Every word needs its own ladder; nobody maintains 300 of them by hand |
| Leitner boxes | Which box the card is in | Open the box, work the pile | The intervals are sound; the failure point is the day you don't open it |
| Algorithmic decks | Your own difficulty ratings | Open the app, clear the queue | Best interval model available — and still gated behind remembering to launch it; a skipped week returns as a backlog |
| Trigger-based review | The algorithm, fired by something you already do | Seconds, unscheduled | Intervals are approximate — they land near the target rather than on it, because the trigger is your day, not a clock |
The last two rows share an interval model and differ only in what starts the review. That difference is where most learning schedules actually die: not in the algorithm, but in the gap between an interval coming due and a person opening an app. A schedule with slightly worse intervals that runs every day beats a perfect one that runs in bursts.
Does the exact interval matter as much as people think?
No — and this is good news, because it means a workable schedule does not require a perfect one. The retention curve in Cepeda et al. (2008) has a broad peak: gaps well to either side of the optimum still produced most of the benefit, and the penalty for being late is much smaller than the penalty for being early. If you have to be wrong, be wrong in the direction of waiting longer.
Two things do matter more than the interval. The first is whether the review is a genuine recall attempt. Karpicke and Roediger (2008) held total study time constant and varied only whether items were restudied or retested; dropping repeated retrieval cut recall a week later from roughly 80% to roughly 35%, while dropping repeated study barely moved it. A perfectly timed glance at a word list is worth less than a badly timed attempt to remember.
The second is whether the review happens at all. Bahrick and colleagues (1993) taught foreign vocabulary across 13 sessions and followed the learners for five years; the 56-day spacing beat the 14-day spacing, but every condition required all thirteen sessions to be completed. That is the assumption almost every published schedule quietly makes and almost no learner meets. Adherence is not a footnote to the spacing effect — over a year, it is the dominant term.
What a real schedule costs per day
People overestimate this badly, and the overestimate is why they never start. A review is not a study session; it is one recall attempt on one word — try to produce the meaning, check, move on. Six seconds is a generous allowance.
A word in active rotation comes back roughly ten times over about six weeks before its intervals stretch past the horizon, which averages to about a quarter of a review per word per day. That single ratio gives you the whole daily bill.
| Words in rotation | Reviews due on an average day | At about six seconds each | What that looks like |
|---|---|---|---|
| 50 | about 12 | just over 1 minute | Less time than one lap of a social feed |
| 100 | about 24 | about 2½ minutes | Comfortably inside the dead time of a single commute |
| 300 | about 72 | about 7 minutes | Still under the length of one queue or one lift ride a day |
| 600 | about 144 | about 14 minutes | The point at which the sessions have to be broken up to survive |
The model: roughly 8–12 spaced encounters per word (Nation & Wang, 1999) spread over about six weeks of expanding intervals, giving ~0.24 reviews per word per day at steady state. New words raise the load, mature words leave it. Treat these as the right order of magnitude rather than a promise — the point is that the daily cost of spacing is minutes, so whatever stops people, it is not the workload.
A schedule that runs without you
Everything above is a scheduling problem wearing a learning problem's clothes. The intervals are known, the daily cost is minutes, and the failure mode is a person forgetting to start. That is the part worth handing to software — and it is the shape LearnScreen is built around.
The interval is decided for you
A Leitner ladder picks the next word: one you just missed comes back almost immediately, one you got right backs off geometrically. You never choose what to review, so the gaps expand the way Cepeda et al. (2008) describe instead of collapsing onto whatever you happen to feel unsure about.
The trigger is already in your day
Phones are checked around 186 times a day, roughly 11.6 times per waking hour (Reviews.org, 2026). LearnScreen uses Apple's Screen Time API to shield the apps you pick, so the card arrives on a reach you were making anyway. Nothing to remember, nothing to launch — the adherence term stops being the weak link.
Six seconds of real recall
The card asks for the meaning before it shows it, then you mark it known or not. That attempt — not the sighting — is what beat rereading in Roediger and Karpicke (2006), and it is the one part of the schedule no system can do on your behalf.
- The dose is yours to set. 3 to 20 words per session, and how often the shield returns is adjustable — typical settings land near 25 recall attempts a day, about two and a half minutes spread across it.
- 1,080+ curated words across 18 topics. Eleven languages to learn, and unlimited custom words with translation, transcription and an example sentence.
- Works offline. Cards and shields run with no network once installed; iCloud backup is opportunistic and writes to your private database, not our servers.
- No account, and no streak to break. Nothing to sign up for, and no chain that punishes a bad day — the queue simply brings that word back the next time you reach for a shielded app.
Keep reading
- How many times do you need to see a word? — the count that these intervals are spreading out.
- How many minutes a day to learn a language? — the time arithmetic, and where the minutes already are.
- How many words a day should you learn? — why the ceiling is review load rather than intake.
- Does passive learning actually work? — what exposure alone produces, and what it does not.
- How many words do you need to speak a language? — how big the rotation has to get in the end.
- Can you learn while doing something else? — which moments in the day can host a review at all.
Frequently asked questions
Sources
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095–1102.
- Bahrick, H. P., Bahrick, L. E., Bahrick, A. S., & Bahrick, P. E. (1993). Maintenance of foreign language vocabulary and the spacing effect. Psychological Science, 4(5), 316–321.
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus' forgetting curve. PLOS ONE, 10(7), e0120644.
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World.
- Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
- Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report
Gap figures are the optimal values reported by Cepeda et al. (2008) for each retention interval they tested, rounded to the nearest sensible unit; the percentages are those gaps expressed as a share of the retention interval. Daily-load figures are arithmetic from the encounter counts in the vocabulary literature, not measured app data, and are given as an order of magnitude.
Let the schedule run itself
LearnScreen keeps the ladder, picks the word and waits for a moment you were having anyway. All it asks for is the six seconds that actually build the memory.
Download on the App Store