Does watching TV in another
language help you learn it?
Yes for listening, much less for new words — and what decides which you get is coverage, not effort. Following ordinary television at 95% takes roughly 3,000 word families; 98% takes closer to 7,000 (Webb & Rodgers, 2009). Below that, an episode is mostly sound you cannot cut into words. What does stick is frequency-gated — how often a word recurred was the strongest predictor of learning it (Peters & Webb, 2018) — and captions beat no captions on both comprehension and vocabulary (Montero Perez et al., 2013).
Every figure on this page is referenced at the end
What watching television has actually been measured to do
Television is the most-recommended and least-quantified piece of language advice there is, which is a shame, because it is one of the better-studied ones. Applied linguists have counted the vocabulary television demands, tested what a single hour of it leaves behind, and meta-analysed whether captions change the answer. The findings are consistent and considerably more specific than “immerse yourself”.
One caveat belongs at the top rather than in a footnote: these studies are short, and they mostly test recognition. A learner watching one documentary under lab conditions is not a learner three seasons into a series, and recognising a word on a test is not the same as producing it in a sentence. What the numbers below are good for is the shape of the effect, not a promise about your December.
| What was measured | What the study found | What it means for you |
|---|---|---|
| The vocabulary television demands | Webb & Rodgers (2009): about 3,000 word families plus proper nouns and marginal words for 95% coverage of television, and closer to 7,000 for 98% | There is a threshold, and it is the single biggest factor. Below roughly 3,000 families, an episode is sound you cannot cut into words |
| What one hour of viewing leaves behind | Peters & Webb (2018): incidental gains from a single hour of L2 television were real but modest, and frequency of occurrence was the strongest predictor of which words were learned | You keep the words the programme happened to repeat. Everything said once is, statistically, a word you met and lost |
| Captions versus no captions | Montero Perez et al. (2013), meta-analysis: captioned video outperformed uncaptioned video on both listening comprehension and vocabulary learning | The single cheapest change available to you is a settings toggle, and it is target-language captions rather than none |
| One series versus assorted shows | Rodgers & Webb (2011): related programmes repeat their vocabulary far more than unrelated ones, because a series keeps its setting, cast and topic | Finishing one show beats sampling ten. Narrow viewing turns single encounters into repeated ones for free |
| Whether the word can be retrieved later | Not measured by viewing studies at all — but Karpicke & Roediger (2008) found retrieval, not re-exposure, is what makes material recoverable | Viewing supplies re-exposure abundantly and retrieval never. That gap is the reason for the familiar ear and the empty mouth |
Read the first row twice. Almost every argument about whether television “works” is really two people at different coverage levels describing two different experiences accurately. The advice is not wrong; it is conditional, and the condition is a number you can estimate: how many words you need to know.
Three claims travel under “just watch TV in Spanish”
They rest on different evidence, which is why the advice feels obviously true and vaguely unsatisfying at the same time. Only the first two are supported, and the second only above a threshold.
- The supported one: your ear improves. This is the real, reliable payoff, and it is not a small one. Extended listening to natural speech trains you to segment a stream into words, to survive speed and accent, and to tolerate not catching everything — a skill no flashcard has ever taught anyone. If your listening is the weak strand, television is close to the ideal instrument.
- The conditional one: this is comprehensible input. True above the coverage threshold and false below it. Webb and Rodgers (2009) put 95% coverage of television at roughly 3,000 word families; comprehension climbs steeply across that band. Below it the material is not input in any useful sense, because you cannot recover meaning from a sentence when you are missing one word in ten and cannot tell where the words end.
- The false one: the words will come by themselves. Pickup is frequency-gated. Peters and Webb (2018) found the number of times a word occurred in the programme predicted whether it was learned, and Nation and Wang (1999) put the encounters needed somewhere between five and twenty. A ninety-minute film hands you most of its interesting vocabulary exactly once — which is precisely the case where the forgetting curve wins.
The general version of this question — what exposure alone delivers and where it stops — has its own page: does passive learning actually work?
Six things decide whether an episode teaches you anything
Two people can watch the same series for a year and end up in completely different places. None of the six reasons is talent, and only one of them is effort. Five are properties of the material and the schedule, which is good news: those are the ones you can change.
| What decides it | What the evidence says | What to do about it |
|---|---|---|
| Your coverage of this programme | Webb & Rodgers (2009): ~3,000 word families for 95% coverage of television, ~7,000 for 98% | Below the threshold, spend some of the time building the frequent core instead — it is what makes the viewing pay |
| How often a word recurs | Peters & Webb (2018): frequency in the programme predicted learning; Nation & Wang (1999): five to twenty encounters to pick a word up | Watch one series to the end rather than ten pilots — narrow viewing multiplies encounters (Rodgers & Webb, 2011) |
| Whether captions are on | Montero Perez et al. (2013): captioned video beat uncaptioned for comprehension and vocabulary; Winke et al. (2010): captions help map sound to written form | Target-language captions as the default. It is one toggle and the best-evidenced change on this list |
| Whether anything is ever retrieved | Karpicke & Roediger (2008): retrieval practice beat repeated study for later recall | Viewing is all confirmation and no recall. Something outside the episode has to ask you |
| The gap between meetings | Cepeda et al. (2006), 254 studies: roughly 47% recall spaced against 37% massed | A plot decides what comes back and when. Nothing in a script is scheduled around what you forgot |
| What happens between episodes | Murre & Dros (2015): the classic forgetting curve replicates — unretrieved material decays fastest right after it is met | The word you looked up on Tuesday needs to come back before Friday, and the next episode will not bring it |
Notice what is missing from that list: attention, discipline and talent. Nothing here fails because you were not concentrating hard enough. It fails because a script has other priorities than your vocabulary — which is exactly what makes it enjoyable enough to keep doing.
The input is real. The second meeting is what’s missing.
Television gets something right that almost every deliberate method gets wrong: it is sustainable. Nobody has to be talked into an episode, nobody breaks a streak, and the input arrives at volume in a form the brain is happy to process for an hour. Nation’s (2007) four strands puts meaning-focused input at roughly a quarter of a balanced programme, and viewing fills that quarter better than anything else available to a learner outside the country. Keep it. The argument here is not that television is a waste of time.
What it cannot do is come back. A series decides what you hear next according to the plot; it has no idea which word you failed on, and no mechanism to reintroduce it three days later when reintroducing it would be worth the most. So the words that recur in the script are learned nearly for free, and the words that occurred once — usually the interesting, mid-frequency ones you actually needed — decay on the ordinary curve (Murre & Dros, 2015). Ten seasons in, the profile is familiar: comfortable comprehension, an ear that copes with speed, and a vocabulary that stalled somewhere around the words the show says every episode.
The fix is small and it is not more viewing. It is a second meeting with the specific words that slipped, spaced out, and shaped as a question rather than a reminder — the pairing Cepeda and colleagues (2006) measured at roughly 47% recall against 37% massed, and the one Karpicke and Roediger (2008) found beats re-reading outright. The awkward part has always been where to put it, because a review session is one more thing to schedule and scheduling is the step people fail at. Which is why it is worth putting on a surface that already recurs by itself: the phone gets picked up around 186 times a day (Reviews.org, 2026), and none of those pickups has to be arranged.
Subtitles: what each setting is actually good for
This is the part of the question people argue about hardest and the part the research answers most cleanly. Montero Perez, Van Den Noortgate and Desmet (2013) meta-analysed captioned video and found it beat uncaptioned video for both listening comprehension and vocabulary learning, and Winke, Gass and Sydorenko (2010) showed why it helps: captions let you map what you are hearing onto a written form, which is the form you will need in order to meet the word again anywhere else.
The four settings are not better and worse so much as good at different jobs. The trap is that the most comfortable one is the least productive.
| Setting | What it is good for | What it costs |
|---|---|---|
| Target-language captions | The best-evidenced default: comprehension and vocabulary both improve over no captions (Montero Perez et al., 2013), and sound gets mapped to written form | Slightly less pure listening practice, and a real temptation to read rather than listen |
| No subtitles at all | Pure segmentation practice — the closest thing to real listening conditions | Below the coverage threshold it produces almost no vocabulary, because an unknown word you cannot even spell is unlookupable |
| Subtitles in your own language | Enjoying a show you would otherwise abandon, and following a complicated plot | The comfortable option that does the least: attention goes to the translation, and the episode becomes reading with a soundtrack |
| A rewatch, captions on the second pass | Turning one episode into two encounters with the same vocabulary — the effect narrow viewing produces across a series (Rodgers & Webb, 2011) | Time. It is the highest-yield option on this table and the one nobody does |
Whichever row you pick, none of them changes the underlying limitation: every setting here decides how well you understand the episode while it is playing. What decides how much you still know in a fortnight is what happens between episodes.
A place for the words the episode said once
That gap between watching and remembering is what LearnScreen was built for. Keep the series. The words you looked up come back on a surface you were already going to look at.
The word you paused on becomes a card
Add it in a few seconds, or paste a whole list after an episode, with translation and an example sentence. Anything you add joins the same queue as the built-in lists — so the vocabulary the script gave you exactly once starts behaving like vocabulary it repeats.
It comes back as a question, not a reminder
Using Apple’s Screen Time API, LearnScreen shields the apps you choose and puts the card where the feed would have been, answer hidden until you commit. That makes each encounter a retrieval attempt — the thing Karpicke and Roediger (2008) found beats re-exposure, and the thing an episode never does.
A Leitner schedule decides what returns
Words you miss come back sooner, words you get right back off geometrically. The spacing is decided by how well you know each word rather than by which episode you happen to be on — the part the distributed-practice research (Cepeda et al., 2006) says does the work.
- Nothing extra gets scheduled. The review lands on pickups you were already making, which is why it survives the weeks when watching an episode is the most you can manage. Defaults come to roughly 25 recall attempts a day — about two and a half minutes, spread out.
- Start from the frequent core if you are under the threshold. 1,080+ curated words across 18 topics in 11 languages, ordered so the cheapest coverage comes first — the fastest route to the 3,000 families that make television legible (which words to learn first).
- Your own words, in bulk. Add a word from the Share sheet or paste a list you built while watching; every entry takes a translation and an example sentence, so the card keeps the context the scene gave it.
- Offline, no account. Cards and shields run without a network once installed, and iCloud backup writes to your own private database rather than our servers.
Related reading
- Does passive learning actually work? — what exposure alone delivers, measured, and where it stops.
- Does listening help you learn a language? — the same question with the picture and the subtitle track taken away.
- How many words do you need to speak a language? — the coverage thresholds that decide whether an episode is input or noise.
- Can you learn while doing something else? — which half of learning survives a divided attention.
- Should you learn words in context or with translations? — what a scene gives a word, and what it withholds.
- Why do I forget vocabulary I’ve already learned? — what happens to a word between episodes.
- How often should you review vocabulary? — the intervals a plot will never supply.
- Should you change your phone’s language? — the same instinct, applied to the other screen in your day.
Frequently asked questions
Sources
- Webb, S., & Rodgers, M. P. H. (2009). Vocabulary demands of television programs. Language Learning, 59(2), 335–366.
- Rodgers, M. P. H., & Webb, S. (2011). Narrow viewing: The vocabulary in related television programs. TESOL Quarterly, 45(4), 689–717.
- Peters, E., & Webb, S. (2018). Incidental vocabulary acquisition through viewing L2 television and factors that affect learning. Studies in Second Language Acquisition, 40(3), 551–577.
- Montero Perez, M., Van Den Noortgate, W., & Desmet, P. (2013). Captioned video for L2 listening and vocabulary learning: A meta-analysis. System, 41(3), 720–739.
- Winke, P., Gass, S., & Sydorenko, T. (2010). The effects of captioning videos used for foreign language listening activities. Language Learning & Technology, 14(1), 65–86.
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
- Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
- Webb, S. (2007). The effects of repetition on vocabulary knowledge. Applied Linguistics, 28(1), 46–65.
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644.
- Nation, I. S. P. (2007). The four strands. Innovation in Language Learning and Teaching, 1(1), 2–13.
- Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report
The honest limits, stated plainly. The viewing studies here are short — single episodes or documentaries, not years of ordinary watching — and they test recognition far more often than production, so they measure the easy end of word knowledge. Coverage figures come from corpus analyses of television, which describe programmes rather than predict any individual’s comprehension. Nothing here was measured on a lock screen: the retrieval, spacing and forgetting findings were established on other materials, and this page applies them to viewing rather than reporting a trial of the combination. Treat it as well-grounded reasoning, and treat any claim that a streaming subscription replaces vocabulary work as having no evidence behind it at all.
Keep watching. Just don’t let the words leave with the credits.
LearnScreen puts the words you looked up on the screen you unlock anyway, hides the answer until you try, and brings back the ones you missed before you forget them.
Download on the App Store