Poster NEAT 2026 — Neuro-AI-Talks Innovatorium, Osnabrück · 14–15 September 2026
Alignment between LMs and brains depends on semantics and syntactic structure
Language models predict brain responses to language remarkably well, but which information carries that alignment is still debated. We erase linguistic concepts from GPT-2's hidden states with LEACE and measure how much brain prediction drops in a naturalistic-listening fMRI dataset (LeBel et al., 9 participants, 27 stories). Erasing part of speech, dependency labels, dependency labels with tree depth, or six GPT-4-labelled semantic features each lowers alignment in the language network. The semantic block costs most, syntactic structure more than word class, and none of it is explained by word rate.
A0 portrait, one page, 2.6 MB, vector text.
Additional plots
The five poster figures at full resolution. All: GPT-2 layer 7, LEACE erasure, cohort mean over 9 participants; language network = Fedorenko SN220 parcels.
-
Figure 1. Held-out encoding accuracy (Pearson r) of the intact GPT-2 layer-7 model, cohort mean on fsaverage; yellow outlines = language parcels. -
Figure 2. Mean held-out r inside the language network per GPT-2 layer (0 = token embeddings); thin lines = participants, thick = cohort mean; shaded column = layer 7, the erasure layer. -
Figure 3. Change in mean held-out r in the language network after erasing each concept (erased minus intact); dots = participants; hatched = label-permuted controls. -
Figure 4. The same change in auditory cortex versus the language network. Solid = concept erased; dotted = the concept with word rate regressed out first; dashed = word rate itself erased. Mean ± SD over participants. -
Figure 5. Where alignment drops: after erasing syntactic structure (blue), after erasing the semantic block (red), and their per-vertex difference (red = semantics costs more, blue = syntax costs more). Shown where the intact cohort mean r exceeds 0.1.
References
Numbers match the superscripts on the poster.
- Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. PNAS, 118(45), e2105646118. doi
- Caucheteux, C., & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology, 5, 134. doi
- Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 36. arXiv
- Abnar, S., Beinborn, L., Choenni, R., & Zuidema, W. (2019). Blackbox meets blackbox: Representational similarity and stability analysis of neural language models and brains. BlackboxNLP @ ACL. arXiv
- Kauf, C., Tuckute, G., Levy, R., Andreas, J., & Fedorenko, E. (2024). Lexical-semantic content, not syntactic structure, is the main contributor to ANN-brain similarity of fMRI responses in the language network. Neurobiology of Language, 5(1), 7–42. doi
- Oota, S. R., Marreddy, M., Gupta, M., & Bapi, R. S. (2023). How does the brain process syntactic structure while listening? Findings of ACL 2023. arXiv
- Reddy, A. J., & Wehbe, L. (2021). Can fMRI reveal the representation of syntactic structure in the brain? NeurIPS 34. proceedings
- Oota, S. R., Gupta, M., & Toneva, M. (2023). Joint processing of linguistic properties in brains and language models. NeurIPS 36. arXiv
- He, L., Zhong, T., Antonello, R., Mischler, G., Goldblum, M., & Mesgarani, N. (2025). Far from the shallow: Brain-predictive reasoning embedding through residual disentanglement. NeurIPS 2025. arXiv
- Hadidi, N., Feghhi, E., Song, B. H., Blank, I. A., & Kao, J. C. (2026). Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds. Nature Communications, 17, 5769. doi
- Toneva, M., Mitchell, T. M., & Wehbe, L. (2022). Combining computational controls with natural text reveals aspects of meaning composition. Nature Computational Science, 2, 745–757. doi
- Merlin, G., & Toneva, M. (2024). Language models and brains align due to more than next-word prediction and word-level information. EMNLP 2024. arXiv
- Oota, S. R., Çelik, E., Deniz, F., & Toneva, M. (2024). Speech language models lack important brain-relevant semantics. ACL 2024. arXiv
- LeBel, A., Wagner, L., Jain, S., Adhikari-Desai, A., Gupta, B., Morgenthal, A., Tang, J., Xu, L., & Huth, A. G. (2023). A natural language fMRI dataset for voxelwise encoding models. Scientific Data, 10, 555. doi
- Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength natural language processing in Python. Zenodo. doi
- Benara, V., Singh, C., Morris, J. X., Antonello, R., Stoica, I., Huth, A. G., & Gao, J. (2024). Crafting interpretable embeddings by asking LLMs questions. NeurIPS 2024. arXiv
- Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., & Biderman, S. (2023). LEACE: Perfect linear concept erasure in closed form. NeurIPS 36. arXiv