/ Research

Why Do I Mishear Lyrics — The Mondegreen Effect

26 September 2026 · CognitionType Research Lab

You have been singing it wrong for years. You know this because someone finally showed you the lyrics on a screen and the words did not match the sounds your brain had been confidently assembling every time the chorus hit. The phrase you heard was vivid, specific, and made its own kind of sense. The actual lyric was something else entirely.

Maybe you spent two decades hearing "wrapped up like a douche" instead of "revved up like a deuce" in Manfred Mann's "Blinded by the Light." Maybe Jimi Hendrix was kissing a guy instead of the sky. Maybe Creedence Clearwater Revival has been directing you to a bathroom on the right instead of warning about a bad moon on the rise since before you were born.

You are not alone. A 2024 analysis found that Elton John is the most mondegreened musician of all time, with more than 2,500 individual reports of misheard lyrics. The phenomenon is so common that it has a name, a seventy-year intellectual history, and a growing body of neuroscience explaining exactly why it happens. What it reveals about how your brain processes language is more interesting than any punchline.

What is a mondegreen and where did the word come from

In November 1954, the American writer Sylvia Wright published an essay in Harper's Magazine called "The Death of Lady Mondegreen." She described a childhood memory of her mother reading aloud from Thomas Percy's Reliques of Ancient English Poetry — specifically, the Scottish ballad "The Bonnie Earl o' Moray." The actual lyrics were:

They hae slain the Earl Amurray,
And laid him on the green.

What young Sylvia heard was:

They hae slain the Earl Amurray,
And Lady Mondegreen.

She preferred her version. Lady Mondegreen was a person, a casualty, a character in a tragedy. "Laid him on the green" was just a detail about grass. Wright's misheard version was emotionally richer, narratively more interesting, and — this is the crucial point — it made enough phonemic sense that her brain never flagged it as wrong.

The word stuck. A mondegreen is now the standard term for a mishearing that substitutes one plausible phrase for another, most commonly in songs but also in prayers, pledges, and any context where sound arrives faster than the brain can parse it with certainty.

Why your brain hears words that are not there

The reason you mishear lyrics is not that your ears are faulty. It is that speech perception was never a passive recording process. Your brain does not receive sound and transcribe it like a dictation machine. It receives a noisy, incomplete acoustic signal and constructs the most plausible interpretation it can, using a combination of what it hears, what it expects, and what it already knows.

Psycholinguists call this top-down processing — the influence of higher-level knowledge on lower-level perception. In 1970, the psychologist Richard M. Warren demonstrated this with an experiment that has become a cornerstone of auditory perception research. He recorded the sentence "The state governors met with their respective legislatures convening in the capital city" and replaced the first "s" in "legislatures" with a cough. Listeners reported hearing the complete word. They could not even identify where the cough had occurred. Their brains filled in the missing phoneme seamlessly, using the surrounding context to reconstruct a sound that was physically absent from the signal.

Warren called this the phonemic restoration effect, and it is one of the most reliable findings in speech perception science. Your brain is so committed to producing coherent language from noisy input that it will hallucinate missing sounds rather than let a gap disrupt comprehension.

The Ganong effect, identified by William F. Ganong in 1980, adds another layer. When listeners hear an ambiguous sound that sits on the boundary between two phonemes — something between a "d" and a "t," for instance — they are more likely to hear whichever phoneme produces a real word. An ambiguous sound at the start of "dash" will be heard as "d." The same sound at the start of "tash" will be heard as "t," because "dash" is a word and "tash" is not. Your lexical knowledge is editing your perception before you are even aware of hearing anything.

These are not bugs in the system. They are features. In a world of background noise, imperfect acoustics, and speakers who mumble, the brain's willingness to fill gaps and resolve ambiguity is what makes spoken language work at all. But when the incoming signal is a song — distorted, reverbed, pitch-shifted, and competing with instruments — those same gap-filling mechanisms start generating creative substitutions.

Why songs are harder to understand than speech

Singing degrades speech intelligibility in ways that spoken conversation does not, and the reasons are acoustic, not cognitive.

When you speak, consonants arrive sharp and fast. The transitions between sounds — the rapid formant shifts that distinguish "b" from "d" from "g" — are preserved because the vocal tract is optimising for clarity. When you sing, the vocal tract optimises for something else: tone, resonance, and pitch. Vowels get elongated to sustain the melody. Consonants get compressed or swallowed because they interrupt the musical line. Research published in the Journal of Voice has shown that vowel intelligibility drops as pitch increases — a singer sustaining a high note is physically distorting the formant structure that makes vowels distinguishable from one another.

The result is an acoustic signal where the phonemic information that matters most for word identification — consonant onsets, formant transitions, the rapid spectral changes that carry meaning — is systematically degraded. Your brain receives fewer cues and must rely more heavily on top-down processing to fill the gaps. When the gaps are large enough and the expectations are strong enough, what you construct from the fragments will sometimes be a different set of words entirely.

Add to this the fact that song lyrics often use poetic compression, unusual syntax, and non-standard vocabulary. You would never say "revved up like a deuce, another runner in the night" in conversation, which means your lexical prediction system has nothing useful to offer. When the bottom-up signal is degraded and the top-down prediction is unconstrained, the system generates its best guess — and its best guess is sometimes a bathroom on the right.

The phonemic processing dimension — why some people mishear more than others

Everyone experiences mondegreens. But not everyone experiences them equally, and the variation is not random.

Research on individual differences in categorical perception — the ability to draw sharp boundaries between similar speech sounds — has shown that typical adults vary meaningfully in how categorically they perceive phonemes. Some listeners weight acoustic cues differently from others. Some have sharper phonemic boundaries. Some are more susceptible to the Ganong effect, meaning their lexical knowledge exerts a stronger pull on their perception of ambiguous sounds.

These differences map directly onto what CognitionType calls the phonemic processing dimension — the ability to perceive, discriminate, and manipulate the sound structure of language. A person with strong phonemic processing can pick out individual sounds in a noisy signal more precisely. They are less likely to mishear a lyric because their bottom-up decoding is delivering cleaner data, leaving less room for top-down construction to fill in something wrong.

A person with weaker phonemic processing — perhaps someone with auditory processing differences, or an unrecognised phonological processing weakness — receives a noisier signal. The acoustic boundaries between phonemes are less crisp. The gap between what arrives and what the brain needs to construct is wider. And wider gaps mean more room for the mondegreen to take hold.

This is why people with dyslexia, which involves a well-documented phonological processing deficit, often report more difficulty following lyrics. The phonological deficit that makes reading effortful — documented extensively by Sally Shaywitz's group at Yale, with effect sizes of around -1.37 standard deviations on phonemic awareness tasks — also makes the acoustic signal of sung language harder to decode. The same dimension is doing the work in both cases. The context is different. The processing challenge is the same.

Memory, sequencing, and why the wrong lyric sticks

Here is the part that frustrates people most: once you have heard the wrong lyric, you cannot unhear it.

This is not stubbornness. It is a property of how memory and sequencing — the second cognitive dimension at play — encodes and retrieves auditory information. When your brain constructs a plausible interpretation of a song lyric and stores it in long-term memory, it does not tag the memory as tentative. It encodes the interpretation as the lyric. Each subsequent listen reinforces that encoding. The auditory trace and the semantic content become fused through repetition, and the phonological loop in working memory — the circuit that rehearses and maintains sound-based information — consolidates the wrong version with the same fidelity it would give the right one.

The Von Restorff effect compounds the problem. This psychological principle, named after the German psychiatrist Hedwig von Restorff, holds that unusual or distinctive information is easier to recall than ordinary information. A misheard lyric that produces something absurd, funny, or vivid — a bathroom, a douche, a guy being kissed — is more distinctive than the intended lyric, which is often poetic and vague. Your brain remembers the mondegreen better than the real lyric because it is more surprising.

Research on earworms and auditory working memory suggests that involuntary musical imagery — songs stuck in your head — preferentially activates the version of the lyric stored in memory, not the acoustic signal itself. If the stored version is the mondegreen, the mondegreen is what loops. Correcting it requires not just learning the right words but overwriting a deeply encoded auditory-semantic association, which is effortful and often incomplete. Months later, in the shower, the old version surfaces because the original trace was never fully erased.

What your eyes do to what your ears hear

There is a reason that seeing the lyrics on screen changes what you hear — and it is not just because you are reading along.

The McGurk effect, discovered by Harry McGurk and John MacDonald in 1976, demonstrated that visual information about mouth movements fundamentally alters auditory perception. When participants watched a video of someone mouthing "ga" while the audio played "ba," they reliably reported hearing "da" — a sound that was neither the visual nor the auditory input, but a fusion of the two. The brain was integrating conflicting sensory channels and producing a compromise percept.

The same principle operates when you watch a music video versus listening to the same song with your eyes closed. Visual cues — a singer's lip movements, on-screen lyrics, even the emotional context of the video — provide top-down constraints that guide phonemic interpretation. This is also why so many people now prefer subtitles: the visual text disambiguates the auditory signal, reducing the cognitive load of phonemic decoding.

When you listen to a song without visual cues — in the car, through headphones, in the background while cooking — you lose that disambiguation channel. Your brain is relying entirely on the degraded acoustic signal plus its own predictions. The mondegreen becomes more likely precisely because one source of corrective information has been removed.

The neuroscience of hearing lyrics that are not there

Brain imaging has begun to map what happens during mondegreen perception at the neural level. A study published in NeuroImage examined brain activity during induced misperceptions of song lyrics using both mondegreens (within-language mishearings) and soramimi (cross-language mishearings, where lyrics in one language sound like words in another). The researchers found that induced misperceptions activated a bilateral network including the middle temporal gyrus and inferior frontal gyrus — regions associated with semantic processing and language production — along with the anterior cingulate cortex, which correlated with how amusing participants found the mishearing.

What this reveals is that mondegreen perception is not a simple decoding error. It recruits the full language network. The brain is actively constructing a meaningful interpretation from an ambiguous signal, integrating semantic knowledge, phonological expectations, and even emotional responses into a coherent percept. The mishearing is not a failure of the auditory system. It is the auditory system working exactly as designed — assembling the most plausible story from incomplete evidence.

The dual-stream model of speech perception, which proposes a left-lateralised dorsal stream for auditory-motor integration and a bilateral ventral stream for extracting meaning from sound, helps explain why mondegreens tend to preserve the phonological shape of the original. The dorsal stream constrains the mishearing to something that sounds similar. The ventral stream then imposes meaning on whatever the dorsal stream delivers. The result is a substitution that is phonemically plausible and semantically creative — which is why mondegreens often feel like they make more sense than the real lyrics.

Attention, rhythm, and the signals your brain misses

There is a third cognitive dimension that shapes mondegreen susceptibility, and it is one that rarely gets discussed in this context: attention and rhythm.

Usha Goswami's research at Cambridge has shown that the brain's ability to track the rhythmic envelope of speech — the rise and fall of amplitude across syllables — is what allows the auditory system to segment the continuous speech stream into meaningful units. When the brain entrains to the rhythm of speech, it knows where to expect the stressed syllables, the word boundaries, and the phonemic transitions that carry the most information.

Singing distorts those rhythmic cues. The natural stress patterns of language get reshaped by the musical meter. A word that would be stressed in speech may fall on an unstressed beat in the song. A syllable boundary that would be clear in conversation gets bridged by a sustained note. The rhythmic scaffold that the attention system normally uses to parse speech into segments is rebuilt according to musical rather than linguistic rules, and the attention system must adapt on the fly.

People with weaker attentional rhythm — less precise neural entrainment to temporal patterns in sound — will find this adaptation harder. They miss the boundaries. The speech stream arrives as a less differentiated flow, with fewer clear segmentation points, and the brain has to carve words out of the stream with less guidance. This is the speech segmentation problem that cognitive scientists have studied for decades: continuous speech has no acoustic equivalent of the spaces between written words, and the brain must infer word boundaries from probabilistic cues. When those cues are distorted by melody and rhythm, the inferences go wrong more often.

What mondegreens tell you about your cognitive profile

The mondegreen is trivial as a punchline and profound as a window into how your brain processes language.

If you mishear lyrics more than most people — if you have spent your life filling in wrong words, struggling to follow vocals against instrumental backing, relying on lyric sheets to learn songs that everyone else seems to absorb by ear — it may not be a quirk or a joke. It may be a signal about your phonemic processing, your auditory working memory, or your attentional rhythm that is worth understanding.

This does not mean there is something wrong with you. It means your brain handles the acoustic-to-linguistic conversion differently from someone else's brain, and the difference is measurable across specific cognitive dimensions. Understanding where you sit on those dimensions — how sharply you discriminate phonemes, how efficiently your working memory holds auditory sequences, how precisely your attention tracks temporal patterns — turns a vague "I'm bad at hearing lyrics" into a specific cognitive profile you can work with.

CognitionType maps your processing style across seven cognitive dimensions, including phonemic processing, memory and sequencing, and attention and rhythm — the three dimensions most directly involved in how your brain decodes sung language. It takes about twelve minutes and returns a profile that helps you understand not just whether you mishear lyrics, but why — and what that pattern reveals about how your mind handles language in every other context, from following conversations in noisy rooms to reading at speed.

The mishearing is the message

The next time you discover you have been singing the wrong words for twenty years, resist the urge to feel embarrassed. The mondegreen is not evidence of carelessness, inattention, or poor hearing. It is evidence that your brain is doing exactly what brains do — constructing meaning from incomplete data, filling gaps with plausible interpretations, and committing to those interpretations with a confidence that makes the world feel coherent even when the signal is noisy.

Every misheard lyric is a small, involuntary experiment in how your particular brain resolves ambiguity. The question is not whether you mishear. Everyone does. The question is what the pattern of your mishearings reveals about the cognitive architecture underneath — the phonemic precision, the memory scaffolding, the rhythmic tracking that together determine how you experience every spoken and sung word that reaches your ears.

Lady Mondegreen was never real. But the cognitive process that created her is the same one that lets you understand speech in a crowded room, follow a lecture while taking notes, and read a word you have never seen before by sounding it out. The brain that invents the wrong lyric is the brain that makes language work. It just occasionally gets creative about it.


CognitionType is an informational assessment, not a clinical diagnosis. If you suspect auditory processing disorder, dyslexia, or another condition affecting how you process spoken language, we encourage you to seek formal evaluation from a qualified professional. A cognitive profile is a complement to clinical assessment, not a replacement.

Discover your own cognitive profile across 7 dimensions.

Take the free assessment