Explainer

Why People Remember Video Over Text (What the Research Actually Says)

The real science on whether people remember video better than text, which popular retention statistics are myths, and how to design training that actually sticks. Solid sources, honestly flagged.
By openCanviz • August 18, 2026

10 min read

You have seen the stat: "people remember 95% of a message from video but only 10% from text." It gets quoted in nearly every pitch for video learning. It is also not from a scientific study. Before you build a training program on a number, it is worth knowing which claims about video and memory are real and which are marketing folklore. The short version: video with narration genuinely helps people learn, the effect is well documented, and the famous percentages people cite for it are mostly made up. This page separates the two.

The myths, named

Three of the most-repeated retention stats do not hold up. Using them costs you credibility with anyone who checks.

  • "95% from video vs 10% from text." This traces to a marketing firm, Insivia, which now states on its own site that the figure came from a small survey of around 200 B2B buyers and was "directional, not a peer-reviewed neuroscience study." It is commonly misattributed to academic sources. Treat it as folklore, not evidence.
  • The "10% of what they read, 20% of what they hear... 90% of what they do" ladder. Often pinned on Edgar Dale's Cone of Experience. Dale never attached any percentages to his cone. The numbers were invented later and have no research basis. They have been repeatedly debunked, including by learning researchers who traced them to sources that do not contain them.
  • "One minute of video equals 1.8 million words." A rhetorical line attributed to Forrester's Dr. James McQuivey, not a measurement. Fine as a figure of speech, not as a statistic.

If a source leans on these numbers, it did not check them. The real case for video is stronger and does not need them.

What the research actually supports

Strip out the folklore and there is a solid, decades-deep body of evidence that words combined with pictures beat words alone. This is the part worth building on.

The multimedia principle

Richard Mayer spent decades running controlled experiments on how people learn from words and pictures. The core finding, the multimedia principle, is that people learn more deeply from words and relevant graphics together than from words alone. In Mayer's own studies the effect was large and consistent. Independent meta-analysis confirms the direction with a more modest effect size, which is the honest picture: the benefit is real, and its size depends on how the material is designed and who is learning.

A related finding, the temporal contiguity principle, is directly relevant to narrated animation: presenting the words and the matching visual at the same moment, rather than separated, produces better learning. A whiteboard animation where the narration and the drawing land together is, in effect, this principle applied.

Dual coding

Underneath Mayer's work is Allan Paivio's dual coding theory: the mind processes verbal and visual information through separate but linked channels. Information encoded in both channels has two routes to retrieval instead of one, which helps it stick. A narrated visual explanation uses both channels by design. Text alone uses one.

The honest caveat

The strongest results compare passive video to passive reading. Active, effortful text study, taking notes, self-testing, explaining it back, can rival or beat passively watching a video. So the accurate claim is not "video always beats reading." It is "well-designed narrated visuals beat plain text for most learners, especially for novel or process material, and active engagement matters more than the medium." That is a claim you can defend.

The part almost everyone skips: spacing

Whichever medium you pick, when people study matters more than most format debates admit. The spacing effect, spreading repetitions over time instead of cramming, is one of the most robust findings in all of learning science, supported by a synthesis of hundreds of experiments. The Ebbinghaus forgetting curve, showing that memory for new material drops steeply then levels off, was formulated in 1885 and successfully replicated in a modern study in 2015.

The practical takeaway is direct: a single long training video, watched once, fights the forgetting curve alone. A set of short lessons revisited over days or weeks works with it. This is the evidence base under microlearning, and it is why breaking training into short, spaced videos outperforms one marathon module for what people remember a month later.

What this means for training you build

The research points to a few concrete design choices, none of which depend on a made-up percentage.

  1. 1

    Pair narration with matching visuals, timed together

    The multimedia and temporal-contiguity effects are strongest when the word and the picture arrive at the same moment, which is exactly what a narrated whiteboard animation does.

  2. 2

    Show processes, do not just describe them

    For a procedure, a visual that builds in order engages the visual channel that plain text leaves idle.

  3. 3

    Keep lessons short and space them

    One idea per short video, revisited over time, works with the forgetting curve instead of against it.

  4. 4

    Add a way to act, not just watch

    Passive viewing is the weakest form of any medium; end a lesson with something the learner does or answers.

  5. 5

    Do not quote the myths

    Make the real case. It is more persuasive to an informed buyer and it will not be embarrassing when someone checks.

Where video specifically helps, and where it does not

Video is not a universal upgrade. It is strongest for novel material, for processes and systems that benefit from being shown in sequence, and for learners who do not already know the subject well. It is weakest as a reference for someone who just needs to look up one fact fast, where a searchable document wins. The mature answer is to match the medium to the moment: narrated visuals to teach something new, text to look something up.

For teaching new processes, that is a strong argument for a narrated, drawn explanation, and it is why a document-first whiteboard workflow fits training so well: it turns the written procedure you already maintain into exactly the narration-plus-matched-visual format the research favors. See how to turn a document into a training video.

Common questions

Is it true that people remember video better than text? For learning new or process-heavy material, well-designed narrated visuals do tend to beat plain text, supported by Mayer's multimedia research and dual coding theory. The specific "95% vs 10%" figure, however, is not from a scientific study and should not be cited as one.

Where does the "95% from video" stat come from? From a marketing firm's small, informal survey, which the firm itself now describes as directional rather than scientific. It is widely repeated but not credible as a research finding.

What actually makes training stick? Pairing words with matching visuals, keeping lessons short and single-topic, spacing them over time, and having the learner do something rather than just watch. The spacing effect and the multimedia principle are the well-supported anchors.

Turn a document into narrated, matched visuals

openCanviz drafts a narrated whiteboard video from your training document, pairing each point with a drawn visual as the research favors. It is free to start. See the case for short, spaced lessons in what is microlearning at /explainers/what-is-microlearning.

Keep reading
Explainer Videos for Compliance Training

How to use explainer videos for compliance training that people actually watch and remember, without rushing the steps a regulator will check. Turn your policy docs into clear, drawn lessons.

How to Make an Employee Onboarding Video

A practical guide to making an employee onboarding video, from what to cover and how long it should be to turning your existing handbook and SOPs into a series of narrated videos.

How to Turn a Document Into a Training Video (Without Animating It by Hand)

You already wrote the lesson plan, SOP, or slide deck. Here is how to turn that document into a narrated training video in five steps, starting from the writing instead of a blank timeline.

8 Best Synthesia Alternatives (2026)

The best Synthesia alternatives in 2026, sorted by what you actually want instead: a different AI avatar tool, a cheaper option, or no avatar at all. Honest picks for training and explainer video.


All Rights Reserved.