Which words should
you learn first?
The most frequent 1,000 word families cover about 72% of written English and over 80% of speech, while the tenth thousand adds under a percentage point (Nation & Waring, 1997; Nation, 2006). Word choice is therefore a bigger lever than study method, and the order is not a matter of taste: finish the first 2,000–3,000 families, then work the mid-frequency bands from 4,000 to 9,000, which Schmitt and Schmitt (2014) show are too rare to absorb from exposure and too common to skip. After that, the best next word is whichever one you just needed and didn't have.
Every figure below is sourced at the end of the page
Word frequency is radically unequal
Most advice treats vocabulary as a pile to be worked through, where any thousand words is worth roughly any other thousand. That is not how language is distributed. Word frequency follows a steep power law — the pattern Zipf (1949) described, in which a small number of words account for most of everything said and written, and the rest thin out into a very long tail.
The practical consequence is that the bands are not interchangeable. Each row below is one thousand word families, in order of frequency, with what that band adds to your coverage of ordinary written text.
| Frequency band | Coverage it reaches | What is in it |
|---|---|---|
| First 1,000 families | About 72% of written text, and over 80% of spoken | Function words and the everyday verbs and nouns every sentence is assembled from — the single highest-value block you will ever learn |
| Second 1,000 | Takes you to roughly 80% | About 8 more points for the same thousand words — still ordinary vocabulary, still nothing you could sensibly skip |
| Third 1,000 | Takes you to roughly 84% | About 4 more points: the last band that pays in whole percentages, and the end of what Schmitt & Schmitt (2014) call high-frequency vocabulary |
| 4th to 9th thousand | Carries you from 84% toward 98% | The mid-frequency range: one to two points per band, too rare to meet often in ordinary input, too common to leave out |
| Beyond the 10th thousand | Under one point per thousand | Zipf's tail — tens of thousands of words, most of which you would meet once in years of reading, and none of which is urgent |
Coverage percentages are for written text and come from Nation and Waring (1997) and Nation (2006); spoken text runs higher at the low bands and lower at the high ones, because conversation reuses a smaller core. The band boundaries are conventions rather than natural joints, and they differ between languages with richer inflection, where a word family covers more distinct forms. What transfers regardless of language is the shape: the curve is steep at the start and almost flat by the tenth thousand.
Three different lists get called “the words to learn”
Arguments about vocabulary selection are usually arguments between people holding different lists and using the same word for all three. They are not rivals, they are stages, and each one is correct in a different place.
- The frequency list. Words ranked by how often they occur in a large corpus. Objectively ordered, impersonal, and unbeatable below the third thousand, because those words appear in everything you will ever read or hear. Its weakness is that it knows nothing about you, which stops mattering only while the words are so common that nobody could avoid them.
- The topic or textbook list. Words grouped by situation — airport, restaurant, office. Useful for a defined near-term need and easy to stay motivated on, but frequency-blind: a themed list will happily teach you a rare noun before a common verb, and semantically related words presented together are known to interfere with each other while they are still new.
- Your own collected words. The ones you looked up because you wanted to understand or say something specific. These carry a real encoding advantage — Laufer and Hulstijn (2001) found that a word you needed, searched for and had to choose a sense for is retained better than one handed to you — but they arrive in no order at all, and left alone they skew rare.
The neighbouring question of how many words the whole job takes, by goal and CEFR level: how many words you need to speak a language.
What makes a word worth your next encounter
Every encounter costs the same few seconds, so the only question that matters is what those seconds buy. Each row is one property that changes a word's value, what the evidence says about it, and how to act on it without keeping a spreadsheet.
| What to check | What the evidence says | What that looks like |
|---|---|---|
| Which band it sits in | Nation & Waring (1997); Nation (2006): the first three thousand families buy coverage in whole points, later ones in fractions | Work the high-frequency core to exhaustion first — no goal changes this order, because it is a fact about the language, not about you |
| Whether input will deliver it | Schmitt & Schmitt (2014); Cobb (2007): mid-frequency words are too rare for reading alone to supply enough encounters | Words from 4,000 to 9,000 are what deliberate review is actually for; the first thousand will be met constantly whatever you do |
| Whether it appears in your input | Webb & Rodgers (2009): about 3,000 families cover 95% of television, 7,000 cover 98% | Above the third thousand, what you actually watch and read is a better selector than a general corpus ranking |
| Whether you needed it yourself | Laufer & Hulstijn (2001): need, search and evaluation raise retention independently of exposure | A word you looked up to say something arrives already half-encoded; keep those, but alongside the core rather than instead of it |
| Whether it is nearly free | de Groot & Keijzer (2000): cognates were learned faster and forgotten more slowly than matched non-cognates | Skim the cognates you can already use — but keep the false friends, which are learned wrongly just as easily |
| How much it brings with it | Bauer & Nation (1993): one family member makes its inflections and regular derivations largely predictable | A base word is worth more than a derived one — learning help hands you helps, helped, helpful, unhelpful |
Only the fourth row is about preference. The other five are about the structure of the language and the structure of your input, which is why selection is a lever you can pull without knowing anything about your own psychology. Note also what is absent: how interesting the word is, how beautiful it sounds, and whether it appeared in the lesson you happen to be on.
The mid-frequency problem
Schmitt and Schmitt (2014) made a proposal that sounds administrative and is not. They divide vocabulary into high frequency — the first 3,000 word families, which every course teaches — low frequency beyond about 9,000, which nobody teaches and nobody needs to, and a mid-frequency range in between that had no name and, largely because of that, no plan. Their argument is that this middle band is where most learners quietly stall: it stands between following a conversation and reading a novel without a dictionary, and it is several times larger than the high-frequency core everyone concentrates on.
What makes the band awkward is a squeeze between two mechanisms. It is too rare for exposure to handle: a word needs something like eight to twelve spaced encounters before it stays (Nation & Wang, 1999; Webb, 2007), and a word in the seventh thousand simply does not appear that often in anything you would read for pleasure. Cobb (2007) modelled exactly this, running large volumes of text through a simulation of incidental learning, and found that reading alone left most of the mid-frequency bands unlearned no matter how much reading was done. It is also too common to skip, since these are the words that make up the difference between 84% coverage and the 98% that Hu and Nation (2000) associate with unassisted comprehension.
That squeeze is the practical case for deliberate, spaced review, and it is oddly specific about where the effort belongs. The first thousand words need almost no help: you will meet them constantly in any input at all. The tail beyond ten thousand needs no help either, because it is not urgent. The mid-frequency range is the one place where the meeting will not happen by itself and the word is worth having — which is to say that a review system's value depends almost entirely on which words you put in it.
Where each source of words runs out
Every way of choosing words works for a while and then stops. Knowing the stopping point is more useful than picking a favourite, because the failure is rarely that the method was wrong — it is that it was kept past the point where it still fitted.
The last row is the one most learners are actually standing on without having noticed.
| Source | What it delivers | Where it runs out |
|---|---|---|
| A general frequency list | The core, in the correct order, faster than any other method | Around the third thousand, where the ranking stops matching your input and the words start feeling arbitrary |
| A course or topic list | A defined situation, and enough structure to keep going | As soon as you need anything outside the topic — and it was never ordered by frequency in the first place |
| Reading and watching | Repeated, contextualised meetings with the high-frequency core, at no cost in effort | In the mid-frequency bands, where Cobb (2007) found the encounters are simply too infrequent to finish the job |
| Words you collect yourself | High involvement load, and perfect relevance to what you were trying to say | Nowhere — but it skews rare and unordered, so it works as a supplement and fails as the whole plan |
The pattern across the four rows is that no single source covers the whole range, and the mid-frequency band is the one that falls between all of them — too high for the beginner list, too low for your own lookups, too rare for input to deliver. What that band needs is not a better source but more spaced encounters than ordinary life supplies: how often to review vocabulary.
Spending the seconds on the right words
If selection is the lever, then a system's job is to put a chosen word in front of you often enough, without asking you to schedule anything. LearnScreen does that with the moment you already have — the reach for a phone — so the effort goes into deciding which words, not into finding time for them.
The list is ordered for you
1,080+ curated words across 18 topics, sequenced so the everyday core comes before the specialised vocabulary rather than in the order a textbook chapter happened to introduce it. Eleven languages are available as learning targets.
Your own words sit alongside it
Unlimited custom words with translation, transcription and example sentence — the mid-frequency and personally-needed words no general list can predict, kept in the same queue as the core instead of in a notebook you stop opening.
The encounters arrive by themselves
Phones get checked about 186 times a day, roughly 11.6 times per waking hour (Reviews.org, 2026). Apple's Screen Time API lets LearnScreen shield the apps you choose and put a word where the feed would have been, so the meetings a rare word needs happen without being scheduled.
- A Leitner schedule decides the order of return. Words you got wrong come back sooner and words you got right back off geometrically, which is what turns a long list into a manageable number of daily encounters.
- You set the dose. Words per session adjust from 3 to 20, as does how often the shield returns — the defaults come to roughly 25 recall attempts a day, about two and a half minutes spread across it.
- The answer stays hidden until you tap. Each card is a recall attempt rather than a reading, which is the direction that transfers: active versus passive vocabulary.
- It works offline, with no account. Cards and shields run without a network once installed; iCloud backup writes to your own private database rather than our servers, and there is nothing to sign up for.
Related reading
- How many words do you need to speak a language? — the size of the whole job, by goal and CEFR level.
- How many words a day should you learn? — the pace the selected list is worked through at.
- Should you learn words in context? — what to put on the card once the word is chosen.
- How often should you review vocabulary? — the intervals a rare word needs to survive.
- Does passive learning actually work? — what exposure alone delivers, and where it stops.
- Why do I forget vocabulary I've learned? — what happens to a word between encounters.
- Should you study vocabulary before bed? — what a night does to the words you chose, and why the last review before sleep is the valuable one.
Frequently asked questions
Sources
- Nation, I. S. P., & Waring, R. (1997). Vocabulary size, text coverage and word lists. In N. Schmitt & M. McCarthy (Eds.), Vocabulary: Description, Acquisition and Pedagogy (pp. 6–19). Cambridge University Press.
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
- Schmitt, N., & Schmitt, D. (2014). A reassessment of frequency and vocabulary size in L2 vocabulary teaching. Language Teaching, 47(4), 484–503.
- Zipf, G. K. (1949). Human Behavior and the Principle of Least Effort. Addison-Wesley.
- Adolphs, S., & Schmitt, N. (2003). Lexical coverage of spoken discourse. Applied Linguistics, 24(4), 425–438.
- Hu, M., & Nation, P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language, 13(1), 403–430.
- Laufer, B., & Ravenhorst-Kalovski, G. C. (2010). Lexical threshold revisited: Lexical text coverage, learners’ vocabulary size and reading comprehension. Reading in a Foreign Language, 22(1), 15–30.
- Webb, S., & Rodgers, M. P. H. (2009). Vocabulary demands of television programs. Language Learning, 59(2), 335–366.
- van Zeeland, H., & Schmitt, N. (2013). Lexical coverage in L1 and L2 listening comprehension: The same or different from reading comprehension? Applied Linguistics, 34(4), 457–479.
- Cobb, T. (2007). Computing the vocabulary demands of L2 reading. Language Learning & Technology, 11(3), 38–63.
- Laufer, B., & Hulstijn, J. (2001). Incidental vocabulary acquisition in a second language: The construct of task-induced involvement. Applied Linguistics, 22(1), 1–26.
- de Groot, A. M. B., & Keijzer, R. (2000). What is hard to learn is easy to forget: The roles of word concreteness, cognate status, and word frequency. Language Learning, 50(1), 1–56.
- Brysbaert, M., Stevens, M., Mandera, P., & Keuleers, E. (2016). How many words do we know? Frontiers in Psychology, 7, 1116.
- Bauer, L., & Nation, P. (1993). Word families. International Journal of Lexicography, 6(4), 253–279.
- Schmitt, N., Jiang, X., & Grabe, W. (2011). The percentage of words known in a text and reading comprehension. Modern Language Journal, 95(1), 26–43.
- Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
- Webb, S. (2007). The effects of repetition on vocabulary knowledge. Applied Linguistics, 28(1), 46–65.
- Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report
The coverage percentages are corpus statistics for English, and the exact figures move with the corpus, the text type and the unit counted — word families, lemmas and word types give different numbers for the same text. Languages with richer inflection or productive compounding distribute differently, so treat the specific percentages as English and the shape of the curve as general. Schmitt, Jiang and Grabe (2011) also found the coverage-comprehension relationship to be broadly linear rather than a cliff, so the 95% and 98% figures are useful landmarks rather than thresholds where understanding switches on. Card counts and daily totals are arithmetic from a roughly twenty-second recall attempt, not measured app data.
Put the right words where you'll see them
LearnScreen keeps a curated core and your own collected words in one spaced queue, and delivers them at the moment you reach for your phone — so the words that ordinary input never repeats often enough get the encounters they need anyway.
Download on the App Store