The short version
Language learning requires encounters with language. Research indicates that audiovisual input can support vocabulary, listening, grammar, pronunciation, speaking, and broader proficiency outcomes, and that captions can often improve comprehension and vocabulary learning. It also shows substantial variation: learners do not absorb everything they encounter, one viewing usually produces modest gains, and the result depends on the person, material, task, support, and measurement.
That is a useful evidence base for spending time with understandable video. It is not evidence that watching anything passively guarantees fluency, that one exact comprehension percentage is optimal, or that every part of language acquisition is explained by a single hypothesis.
Start with a distinction: theory versus evidence
A theory organizes explanations and predictions. An individual study tests a bounded question with particular participants and materials. A systematic review or meta-analysis combines studies using explicit methods, but its result is still limited by the studies available and the comparability of their designs.
The Input Hypothesis is historically important because it placed understandable, meaning-bearing language at the center of acquisition. The shorthand i+1 gave teachers and learners a memorable way to think about manageable novelty. Other research traditions emphasize interaction, output, attention, explicit knowledge, skill practice, memory, social participation, and motivation.
These perspectives are not interchangeable, and the field has not reduced acquisition to a single settled mechanism. InputScout therefore uses comprehensible input as a practical design orientation—not as a claim that one hypothesis has been proven in its strongest form.
What audiovisual-input studies suggest
Audiovisual input combines spoken language with a moving visual context. The image can clarify who or what is being discussed, show an action, establish a setting, and convey emotion or intention. Those cues may reduce ambiguity and help learners maintain a representation of the message.
Sutton and Webb’s 2026 meta-analysis synthesized 56 experiments on audiovisual input without on-screen text. It reported learning across measured language domains and found variation by video category, with educational material producing different outcomes from entertainment-focused material. The authors used a largely within-group pre/post evidence base, which is important when interpreting effect sizes because improvement cannot always be separated cleanly from testing or other influences.
Montero Perez’s 2022 review surveys a broader body of audiovisual research, including comprehension, vocabulary demands, on-screen text, and learner perceptions. The review supports the potential of viewing while emphasizing that proficiency, vocabulary knowledge, caption mode, genre, and study design complicate simple prescriptions.
The careful conclusion is that video is a credible source of language-learning opportunities. The medium itself does not determine whether a particular learner will understand or learn from a particular clip.
Captions and subtitles
Researchers commonly distinguish captions—on-screen text in the same language as the audio—from subtitles translated into another language. Everyday platform labels are less consistent.
The 2013 meta-analysis by Montero Perez and colleagues found an overall advantage for captioned video in listening comprehension and vocabulary outcomes across the included studies. A newer meta-analysis by Kurokawa, Hein, and Uchihara synthesized a larger captioned-viewing literature and found a medium overall advantage for incidental vocabulary learning over uncaptioned viewing, while also identifying moderators and methodological caveats.
Captions may help learners segment continuous speech, recognize word forms, and connect audio with spelling. They can also redirect attention toward reading, and benefits depend on what the learner is trying to accomplish. Translated subtitles may be the support that makes a difficult story accessible, but they can make it easier to follow the translation instead of processing the target-language audio.
The practical response is not “captions always on” or “captions are cheating.” Use the least support that lets you engage successfully with the chosen material, and change the support when the goal changes.
Incidental vocabulary learning
Incidental learning means vocabulary develops while attention is primarily on understanding a message rather than preparing for a vocabulary test. Studies show that this can happen through viewing. They also show that gains from a single exposure tend to be selective.
Peters and Webb’s television study, for example, examined word learning after viewing and considered prior vocabulary knowledge and frequency of occurrence. Across the broader literature, repeated encounters, available context, existing knowledge, attention, and the way knowledge is tested all matter.
Recognition often develops before a learner can recall and use a word. Knowing that a form was present is different from explaining it, understanding all its senses, pronouncing it, combining it naturally with other words, or retrieving it in conversation. Claims about “words learned” should therefore be read alongside the test used.
This is why InputScout does not equate minutes watched with vocabulary acquired. Viewing creates opportunities. Learning from those opportunities is gradual and variable.
Repetition and extensive practice
Repeated encounters give a learner more chances to retrieve an earlier interpretation, notice another feature, and strengthen a form–meaning connection. Rewatching can also reduce the comprehension burden because the storyline is already known.
Research on extensive reading offers related evidence for sustained, meaning-focused exposure over time. Nakanishi’s meta-analysis found positive effects from extensive reading, though reading and audiovisual viewing are not identical activities. The transfer is conceptual: accessible volume and continuity can matter, and a single difficult encounter is a poor model for a long-term practice.
Repetition is not automatically productive. Attention can disappear when material becomes overfamiliar, and repeating one clip forever limits the range of language encountered. A useful practice mixes strategic rewatching with new, related material.
Comprehension thresholds and lexical coverage
Researchers have used lexical coverage to examine how knowing different proportions of running words relates to comprehension. Higher coverage generally helps. Exact thresholds depend on the medium, task, text, learner, and definition of adequate comprehension.
Vocabulary coverage cannot fully describe video difficulty. Speech rate, accent, reductions, syntax, discourse organization, audio quality, visuals, topic knowledge, humor, and narrative structure all contribute. A threshold found for reading a particular kind of text should not be presented as a universal law for watching every video.
InputScout therefore asks learners for direct fit feedback and shows difficulty confidence. It does not claim that its score converts to a percentage of words known or understood.
Attention, interaction, and output
Understanding a message can expose patterns without requiring conscious analysis of every form. At the same time, attention influences what is encoded, and learners sometimes benefit from deliberately noticing a recurring expression or contrast.
Interaction adds opportunities to request clarification, confirm an interpretation, and receive feedback. Output—speaking or writing—requires retrieval and exposes gaps that receptive activity may conceal. Explicit explanation can make a pattern easier to see, and deliberate practice can strengthen access to important forms.
These are reasons to treat viewing as a foundation or strand of a broader practice rather than an ideological replacement for everything else. If your priority is spontaneous conversation, include conversation. If a grammar issue repeatedly blocks comprehension, investigate it. If a word is urgent and important, study it deliberately.
Motivation is not a decorative extra
Interest affects whether a learner pays attention, finishes a video, returns tomorrow, and builds enough volume for repeated encounters. This does not mean that enjoyable content magically produces acquisition. It means a theoretically ideal resource has limited value if the learner consistently avoids it.
InputScout includes topic and dialect preferences alongside difficulty because sustainable exposure depends on both fit and desire. It does not use engagement popularity as a substitute for either.
How to read a language-learning claim
When an article promises a result, ask:
- Was the evidence a theory, an observational association, a controlled experiment, or a synthesis?
- Who were the learners, and what languages and proficiency levels were involved?
- What exactly did they do, for how long, and with what support?
- Was the outcome immediate recognition, delayed recall, comprehension, general proficiency, or self-report?
- Was there a comparison group, and were assessors or participants aware of the condition?
- Are the practical recommendations narrower than the headline?
InputScout’s research pages link to DOI or publisher records so readers can inspect the source rather than relying on a slogan.
What InputScout concludes
The evidence supports offering learners more accessible, interesting audiovisual input; making captions and filters visible; encouraging repeated and sustained practice; and acknowledging uncertainty. It also supports modest language: results vary, scores are estimates, and viewing works best as part of a practice shaped by the learner’s goals.
That is enough to build a useful product without pretending the science is simpler than it is.