Why do humans and LLMs encode linguistic structure?
In classic linguistic theory, parsing a sentence into hierarchical constituents is considered a fundamental step for comprehension. Yet, while language users are largely unaware of this parsing process, modern large language models (LLMs), which achieve human-level performance on many tasks, also lack explicit representations of linguistic constituents. This talk will explore this apparent paradox. First, I will present recent evidence demonstrating that the objective of next-word prediction can force LLMs to implicitly learn sophisticated constituency analysis. I will detail a series of experiments using a novel one-shot learning paradigm to test and compare the behavioral sensitivity to linguistic constituents in both human participants and LLMs. Second, I will examine a potential trade-off between the precision and efficiency of next-word prediction. Our experiments indicate that the human brain may optimize this process by constructing more compact, structure-based contextual representations, thereby achieving greater predictive efficiency. Collectively, this research suggests that next-word prediction drives the implicit acquisition of linguistic structure, which in turn facilitate more effective word prediction.