Scientists examined whether human anticipatory language processing works the same way as next-word prediction in large language models (LLMs). The new study, published in Nature Neuroscience, reports that the brain predicts upcoming language using larger grammatical units, or constituents, rather than relying solely on the single most likely next word. "While LLMs are trained and optimized to predict the next word, the human brain makes predictions by grammatically grouping words into phrases," says coauthor David Poeppel.
The researchers ran several experiments with Mandarin Chinese speakers and recorded neural responses with magnetoencephalography (MEG). They complemented those recordings with behavioral Cloze tests, where readers supply missing words, and they reanalyzed additional brain data from patients exposed to English to test language generality.
- They used LLMs to quantify predictability via entropy and surprisal.
- High entropy means many possible next words; high surprisal means a word is unexpected.
- They compared model-based predictions to time-locked brain responses for the same sentences.
- Instead of a uniform match, brain correlations depended on a word’s position in grammatical phrases.
These results indicate that human prediction is modulated by grammatically organized chunks, a sensitivity that LLM next-word probabilities do not capture. The findings raise questions about how computational models should represent linguistic structure to better reflect human language processing.
Difficult words
- anticipatory — predicting something before it happens
- constituent — a grammatical unit like a phraseconstituents
- magnetoencephalography — a brain activity recording method
- Cloze test — task where readers fill missing wordsCloze tests
- entropy — measure of how many choices exist
- surprisal — measure of how unexpected something is
- modulate — change or influence the strength or levelmodulated
Tip: hover, focus or tap highlighted words in the article to see quick definitions while you read or listen.
Discussion questions
- What benefits might there be if computational models represented grammatical chunks like the human brain?
- How does using both Mandarin and English data affect the study's claims about general language processing?
- How could these findings influence the design of text-prediction features in real applications?
Related articles
Researchers find 'vibe coding' linked to insecure AI-written code
A research team found that a programming style called "vibe coding" is producing insecure code with help from generative AI tools. A new radar scans public vulnerability data and flags cases that show AI signatures or risky patterns.
Zenica School of Comics: Art and Education for Children
The Zenica School of Comics began during the 1992–95 war and has taught around 200 young artists. The school still runs, faces changes from tablets and AI, and the regional comics scene survives through festivals and cooperation.
How civil society adapts to AI and surveillance
In April 2026 Global Voices and IRIS published ten case studies from the Global Majority showing how civil society groups respond to AI and algorithmic platforms. Responses include co‑opting, countering and innovating, with local and cross‑border strategies.