You can follow the conversation. You catch the jokes. But when it is your turn, the word you need arrives a beat too late, and the moment is gone. The reason is simple and fixable: understanding and speaking run on two different skills. Recognizing a word when you hear or read it is one skill. Pulling that word out of memory on demand, fast enough for a live sentence, is a separate skill called retrieval, and most studying never trains it. Conversation runs almost entirely on retrieval, and retrieval is trainable on its own, without a conversation partner, in minutes a day.
Updated August 2026: rewritten around the research on retrieval practice, typing versus speaking, and word combinations, with sources linked where each claim is made.
Recognition is not retrieval
Every learner’s passive vocabulary (words you understand) runs far ahead of their active vocabulary (words you can produce). That lag is not a personal failing; it is the documented norm in second-language research (Laufer, 1998, Language Learning). Understanding gives you generous time: the word is right there in the sentence, and context does half the work. Speaking gives you no time at all.
A half-known word costs you a beat. By the time it arrives, the next sentence is already gone.
So the gap you feel in Spanish is not evidence that you learned the wrong words. It is evidence that you trained recognition and never trained retrieval.
What actually trains retrieval
The research here is unusually one-sided. Actively recalling a word from memory beats re-reading or re-viewing it by about half a standard deviation of extra retention, measured across 159 separate comparisons (Rowland, 2014, Psychological Bulletin; the effect size is g = 0.50, where g is a standard unit for sizing the difference between two methods: 0.2 counts as small, 0.5 as medium, 0.8 as large, so this is a solid, clearly noticeable advantage). The classic experiment used foreign-language vocabulary specifically: repeated self-testing produced large gains a week later, repeated studying produced almost none, and learners could not feel the difference while it was happening (Karpicke and Roediger, 2008, Science).
Three details in that literature matter for the speaking problem:
- Direction matters. Practicing from the meaning to the word builds the ability to produce the word; going only from the word to the meaning does not transfer well to production (Terai, Yamashita and Pasich, 2021, Studies in Second Language Acquisition).
- Retrieval practice builds speed, not just accuracy. Words you have actively recalled come back faster next time (van den Broek and colleagues, 2014, Memory), and speed is the point: fast, automatic access to vocabulary predicts speaking fluency, while the sheer size of your vocabulary does not (Takizawa and colleagues, 2025, Applied Linguistics, a study of 210 learners).
- Spacing beats cramming. Spreading retrievals out over days beats massing them into one session, across 48 experiments (Kim and Webb, 2022, Language Learning).
Do you have to speak to close the gap?
Honest answer: partly, and it helps to know exactly which part.
The words themselves transfer. In the cleanest direct experiment, learners who studied words by writing scored 68 percent when tested in writing and 58 percent when tested in speech, and even that gap disappeared within a week; the researchers went in expecting words to work best in the mode they were learned in, and their data rejected that prediction (Candry, Deconinck and Eyckmans, 2018, Journal of the European Second Language Association). Typed recall genuinely builds the words you will later speak.
What typing cannot do is automatize the act of speaking itself. Skills automatize separately: comprehension practice speeds up comprehension, production practice speeds up production, and neither substitutes for the other (DeKeyser, 1997). Speaking fluency in training studies improves through spoken repetition under time pressure, such as telling the same story in four, then three, then two minutes (de Jong and Perfetti, 2011, Language Learning).
The clean way to say it: typed retrieval builds and speeds the engine; speaking connects that engine to your mouth. A typed quiz alone will not make you a fluent speaker, and you should distrust any app that claims otherwise. The bridge costs nothing: say your answer out loud as you type it (saying items aloud adds a reliable memory edge over silent study, shown durable for foreign vocabulary: Icht and Mama, 2019, Language Teaching Research), and retell things you can already say, a little faster each time.
Learn combinations, not only words
In Spanish you do not really “take” a walk, you give one: dar un paseo. Knowing paseo is not the same as being able to say dar un paseo when the moment comes.
The research on multi-word combinations is honest in both directions. They are not easier to memorize; in direct comparisons, combinations carried a heavier learning burden than single words (Peters, 2016, Language Teaching Research). But they are what fluent speech is made of: natural-sounding speakers assemble sentences largely from prefabricated chunks (Pawley and Syder, 1983), and learners’ chunk knowledge lags badly, sitting below 50 percent even for combinations built from the 1,000 most frequent word families (Nguyen and Webb, 2017, Language Teaching Research). Deliberately studying combinations closes that gap, with large gains in a 2024 meta-analysis of 17 studies (Li and Lei, 2024, IRAL; d = 1.415, where d is another effect-size unit and anything past 0.8 counts as large, though this one pools varied study designs, so read it as “works well”, not as an exact constant).
There is even a memory trick hiding in the combination: attaching a new word to a word you already know gives it a retrieval cue you will actually meet in real sentences (Kasahara, 2011, System). So when a word arrives inside a combination, save the combination.
What this looks like day to day
The routine the evidence supports is small and specific:
- See the meaning, produce the word from memory. Meaning to word, not the reverse.
- Say it out loud while you type it.
- Space the repetitions out; let the misses come back sooner than the hits.
- Watch your speed improve, not just your accuracy.
- When a word shows up inside a combination, keep the combination.
This pattern is what Koboshi’s Hard quiz is built on: typed recall from the meaning, spaced by a scheduler, and, since August 2026, timed, because response speed is the number the fluency research says actually matters. Phrases and multi-word expressions can be saved alongside single words; a dedicated combination exercise is something we may add in the future.
Whichever tool you use, the shape of the fix is the same: a few minutes of producing words from memory, daily, beats hours of re-reading, weekly. The gap between understanding and speaking is normal, it is measurable, and it closes from the retrieval side. The words stop arriving late.
For years, when someone has asked me whether I speak Spanish, my answer has not been “hablo un poco de español”. It has been “entiendo más de lo que puedo hablar”: I understand more than I can speak. Every learner carries a version of that sentence. Train the retrieval side, and someday we will not need to say it.