Why do humans and LLMs encode linguistic structure?

In classic linguistic theory, parsing a sentence into hierarchical constituents is considered a fundamental step for comprehension. Yet, while language users are largely unaware of this parsing process, modern large language models (LLMs), which achieve human-level performance on many tasks, also lack explicit representations of linguistic constituents. This talk will explore this apparent paradox. First, I will present recent evidence demonstrating that the objective of next-word prediction can force LLMs to implicitly learn sophisticated constituency analysis. I will detail a series of experiments using a novel one-shot learning paradigm to test and compare the behavioral sensitivity to linguistic constituents in both human participants and LLMs. Second, I will examine a potential trade-off between the precision and efficiency of next-word prediction. Our experiments indicate that the human brain may optimize this process by constructing more compact, structure-based contextual representations, thereby achieving greater predictive efficiency. Collectively, this research suggests that next-word prediction drives the implicit acquisition of linguistic structure, which in turn facilitate more effective word prediction.

Next
Next

When does subtitle viewing become reading? Identifying the onset of reading in dynamic multimodal environments.