Do language learning apps
actually work?
Yes, measurably, and mostly at the beginner end. About 34 hours of app study matched a first university semester of Spanish on a placement test — in a trial the app’s own maker commissioned (Vesselinov & Grego, 2012). An independent semester with the same app found real gains in reading and listening and almost none in speaking (Loewen et al., 2019). The variable that separates learners it works for from learners it doesn’t is not the app: it is how many sessions happen.
Every figure on this page is referenced at the end
What the efficacy studies actually found
Language apps are one of the most downloaded categories on earth and one of the least rigorously evaluated. There are a handful of real studies, they are small, and their results are more interesting than either the advertising or the backlash. Read them together and a consistent shape appears: genuine gains, concentrated in recognition rather than production, with enormous variation between individuals.
The funding column matters as much as the findings column. The single most-quoted number in this space comes from a study the vendor paid for, which does not make it wrong but does make it a floor for scepticism rather than a ceiling for claims.
| Study | What it measured | What it found |
|---|---|---|
| Vesselinov & Grego (2012) | Adult beginners in Spanish, tested on a university placement test before and after self-paced app study — commissioned by the app’s maker | An average of roughly 34 hours produced placement gains comparable to a first university semester, with very wide variation in how much each learner actually studied |
| Loewen et al. (2019) | Independent semester-long study of university learners using the same app to study Turkish from scratch | Measurable gains in reading and listening, very little movement in speaking, and steep drop-off in use across the term |
| Rachels & Rockinson-Szapkiw (2018) | A gamified Spanish app compared against face-to-face instruction in elementary classrooms, achievement and self-efficacy both measured | No significant difference in achievement, and lower self-efficacy in the app group — equal results, less confidence |
| Golonka et al. (2014) | Review of the whole technology landscape for language learning, sorted by technology type and strength of supporting evidence | Plenty of enthusiasm, thin rigorous evidence that any given technology beats ordinary instruction; the best-supported uses were narrow and specific |
| Burston (2015) | Two decades of mobile-assisted language learning implementations, analysed for what the reported learning outcomes actually rest on | Most projects were very short and very small, so positive results describe weeks of novelty rather than terms of learning |
None of these studies tested this app, and none of them measured fluency. They measure placement scores, receptive skills and classroom achievement over weeks or a term. Treat the category verdict as “does something real, at the beginner end, in recognition first” — and treat any number promising conversational ability from screen taps as unsupported.
Three different claims travel under “apps work”
They rest on completely different evidence, which is why the same person can reasonably believe apps are useful and that app marketing is nonsense. Only the first one is well supported.
- The supported one: measurable beginner gains, mostly receptive. Every study above found something. Placement scores move, reading and listening move, and vocabulary is the part that moves most reliably — it is the component that responds best to spaced retrieval, which is the one thing software does better than a classroom.
- The overstated one: an app replaces a course. Rachels and Rockinson-Szapkiw (2018) found parity with classroom teaching on achievement, not superiority, and Loewen and colleagues found the speaking gap the classroom exists to close. Nation’s (2007) four strands make the shape explicit: input, output, deliberate study and fluency development are separate jobs, and a tap-and-select app is strong at exactly one of them.
- The false one: fluent in fifteen minutes a day. No study in this literature measured fluency at all, so no study supports a fluency claim. At ten minutes a day, the 34 hours in the best-known result takes about seven months — and that arithmetic assumes you never skip, which is precisely the assumption the next section takes apart.
Exposure without retrieval is a fourth claim again, with its own evidence base: does passive learning actually work?
What actually decides whether an app works for you
Six things, each taken from a study rather than from a feature list, and ordered by how much they move the result. The first one dwarfs the rest, which is inconvenient for everybody selling the other five.
| What decides it | What the evidence says | What that looks like |
|---|---|---|
| Whether the session happens | Nielson (2011): federal employees given licences, paid study time and tutor support still largely stopped early, and few reached meaningful usage | The failure mode is not a bad lesson, it is a lesson that never opened — if support and paid time were enough, that study would have ended differently |
| Whether you retrieve or recognise | Roediger & Karpicke (2006): being tested on material produced better long-term retention than restudying it for the same time | Picking the right tile out of four is much easier than producing the word, and easier practice buys less; the exercise type matters more than the app |
| Whether practice is spread out | Cepeda et al. (2006), synthesising 254 studies: about 47% recall for spaced practice against 37% for the same practice massed | Ten minutes on six days beats an hour on Sunday, and the app that shows up daily wins on this axis before any content is compared |
| Whether each word gets enough meetings | Nation & Wang (1999): a word generally needs somewhere around 8 to 12 spaced encounters before it stays | A streak made of new material every day quietly never finishes anything; the words that stick are the ones scheduled to come back |
| Whether the words are worth the meetings | Nation (2006): roughly the first 3,000 word families cover about 95% of everyday conversation, and coverage flattens sharply after that | Frequency order is most of the return in the first year, which is why a themed course of animal names feels productive and measures poorly |
| Where the minutes come from | Reviews.org (2026): US adults report checking their phones about 186 times a day, roughly 11.6 times per waking hour | A plan that needs a new slot in the day competes with everything else in it; one that rides an existing habit does not |
Only rows two to five are about software quality, and they are the rows every app in the category has largely solved. Rows one and six are about delivery, and they are where the outcomes are actually decided — which is the finding that keeps making method comparisons come out flat.
The variable the reviews never measure
App comparisons are written about features: how good the speech recognition is, whether the grammar notes are any good, how the review queue is scheduled. Those are real differences and they are small ones. The difference that is not small is between a learner who did forty sessions and a learner who did four, and no review can tell you which of those you will be, because it is not a property of the app.
This is why the efficacy literature keeps producing flat results. Golonka and colleagues (2014) went looking for technologies with strong evidence of advantage and mostly found enthusiasm; Burston (2015) found that two decades of mobile-assisted projects were dominated by studies lasting weeks. When time-on-task varies by an order of magnitude between participants and the trial is short, a method difference has almost no room to show up. Meanwhile Nielson (2011) ran what should have been the best case for self-study software — volunteers, licences, protected time, tutor support — and watched most of them stop. Adherence is not a soft factor sitting alongside the hard ones. It is the hard one.
Which reframes the question people actually mean when they ask whether apps work. The honest version is: does this app produce enough sessions, on days when nothing is going well, for the spacing and retrieval effects to accumulate? Every technique on this page is worthless in an app you stopped opening in March, and even a mediocre technique compounds in one you never had to decide to open. The productive move is not to hunt for a better exercise. It is to attack the step where the losses actually are.
What quitting costs, in the units the research uses
“Be consistent” is advice nobody can act on. What follows is the same point with numbers attached: four specific things that stop happening when the sessions stop, each of them measured by somebody.
Notice that three of the four are not about learning less. They are about losing material you had already paid for.
| What is lost | Evidence | Measured size |
|---|---|---|
| The spacing advantage | Cepeda et al. (2006), synthesising 254 studies of distributed practice | About 47% spaced against 37% massed — a gap you only collect by returning on a later day |
| The encounters a word still needs | Nation & Wang (1999), on how often a word must be met before it holds | Roughly 8 to 12 spaced meetings, so a word abandoned at four is not four-twelfths learned — it is gone |
| Everything already half-learned | Murre & Dros (2015), replicating Ebbinghaus’ forgetting curve | Steep early loss without review — the queue you stop clearing decays fastest in the first days, before it slows |
| The finished course | Nielson (2011), workplace self-study with licences, time and support | Most participants stopped early — the modal outcome of self-directed software is an unfinished course, not a bad one |
This is not an argument for discipline. It is an argument for arranging things so that fewer decisions stand between you and a review — which is a design problem, and a solvable one. The habit side of it is on how to stop quitting a language.
Removing the step where the losses happen
If adherence is the variable that decides the outcome, the useful design question is not how to make a lesson better. It is how to make a review happen without anyone deciding to start one. LearnScreen is built around that single idea — here is the mechanism, plainly.
No session to start
Apple’s Screen Time API lets the app shield the apps you choose. When you open one, a word card appears where the feed would have been. Phones are checked around 186 times a day (Reviews.org, 2026), so the review rides a habit that already exists instead of asking for a new slot in the day.
Each card is a retrieval attempt
The answer stays hidden until you tap, so the card is a test rather than a reading — the direction Roediger and Karpicke (2006) found produced better long-term retention. Ambient delivery raises the number of attempts; it does not change what one attempt is worth.
A schedule decides what returns
A Leitner queue brings missed words back sooner and lets known words back off geometrically, so the 8 to 12 meetings a word needs (Nation & Wang, 1999) get spread across days rather than crammed into one sitting.
- You set the dose. Words per session adjust from 3 to 20, as does how often the shield returns. The defaults come to roughly 25 recall attempts a day — about two and a half minutes spread across it, which is small enough that no other plan has to be cancelled to fit it.
- It is not a course, and does not pretend to be one. This is deliberate vocabulary study, one of Nation’s (2007) four strands. Speaking practice, reading and listening are jobs it does not do, and the honest place for it is underneath them: active versus passive vocabulary.
- Frequency-ordered lists, or your own words. Curated sets start where coverage is cheapest (Nation, 2006), and anything you add by hand or paste in bulk joins the same queue — which words to learn first.
- It works offline, with no account. Cards and shields run without a network once installed; iCloud backup writes to your own private database rather than our servers, and there is nothing to sign up for.
Related reading
- Does passive learning actually work? — what exposure alone delivers, and where it stops.
- How do you stop quitting a language? — the adherence problem this page ends on, taken seriously.
- How many minutes a day do you need? — what those 34 hours look like spread across a year.
- How often should you review vocabulary? — the schedule that turns sessions into retention.
- Active vs passive vocabulary — why app gains land in recognition before production.
- Which words should you learn first? — making each of those encounters worth having.
Frequently asked questions
Sources
- Vesselinov, R., & Grego, J. (2012). Duolingo Effectiveness Study. City University of New York / University of South Carolina. Commissioned by Duolingo.
- Loewen, S., Crowther, D., Isbell, D. R., Kim, K. M., Maloney, J., Miller, Z. F., & Rawal, H. (2019). Mobile-assisted language learning: A Duolingo case study. ReCALL, 31(3), 293–311.
- Rachels, J. R., & Rockinson-Szapkiw, A. J. (2018). The effects of a mobile gamification app on elementary students’ Spanish achievement and self-efficacy. Computer Assisted Language Learning, 31(1–2), 72–89.
- Nielson, K. B. (2011). Self-study with language learning software in the workplace: What happens? Language Learning & Technology, 15(3), 110–129.
- Golonka, E. M., Bowles, A. R., Frank, V. M., Richardson, D. L., & Freynik, S. (2014). Technologies for foreign language learning: A review of technology types and their effectiveness. Computer Assisted Language Learning, 27(1), 70–105.
- Burston, J. (2015). Twenty years of MALL project implementation: A meta-analysis of learning outcomes. ReCALL, 27(1), 4–20.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
- Nation, I. S. P. (2007). The four strands. Innovation in Language Learning and Teaching, 1(1), 2–13.
- Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644.
- Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report
These are small studies. Sample sizes run from single classrooms to low hundreds, durations from weeks to one term, and outcome measures are placement tests and skill batteries rather than real-world conversation. One of them was paid for by the company it evaluated, which is stated where it is quoted. LearnScreen has not been through an efficacy trial and this page does not claim otherwise; the app-specific figures here are arithmetic from its default settings and a roughly twenty-second recall attempt, not measured user data.
Fix the step that decides the outcome
If the research says adherence beats method, the fix is not a better lesson — it is a review that happens without a decision. LearnScreen puts one on the phone you were already unlocking, and asks for no new minutes at all.
Download on the App Store