Poster NEAT 2026 — Neuro-AI-Talks Innovatorium, Osnabrück · 14–15 September 2026

Alignment between LMs and brains depends on semantics and syntactic structure

Abraham Jacob (Bram) Fresen1 & Micha Heilbron1,2
1 Language and Predictive Computation Group, Max Planck Institute for Psycholinguistics, Nijmegen · 2 Amsterdam Brain and Cognition, University of Amsterdam

Language models predict brain responses to language remarkably well, but which information carries that alignment is still debated. We erase linguistic concepts from GPT-2's hidden states with LEACE and measure how much brain prediction drops in a naturalistic-listening fMRI dataset (LeBel et al., 9 participants, 27 stories). Erasing part of speech, dependency labels, dependency labels with tree depth, or six GPT-4-labelled semantic features each lowers alignment in the language network. The semantic block costs most, syntactic structure more than word class, and none of it is explained by word rate.

Poster: background, methods, five result figures, a checks table and conclusions on GPT-2 concept erasure versus fMRI.
Download poster (PDF)
A0 portrait, one page, 2.6 MB, vector text.

Additional plots

The five poster figures at full resolution. All: GPT-2 layer 7, LEACE erasure, cohort mean over 9 participants; language network = Fedorenko SN220 parcels.

References

Numbers match the superscripts on the poster.

  1. Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. PNAS, 118(45), e2105646118. doi
  2. Caucheteux, C., & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology, 5, 134. doi
  3. Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 36. arXiv
  4. Abnar, S., Beinborn, L., Choenni, R., & Zuidema, W. (2019). Blackbox meets blackbox: Representational similarity and stability analysis of neural language models and brains. BlackboxNLP @ ACL. arXiv
  5. Kauf, C., Tuckute, G., Levy, R., Andreas, J., & Fedorenko, E. (2024). Lexical-semantic content, not syntactic structure, is the main contributor to ANN-brain similarity of fMRI responses in the language network. Neurobiology of Language, 5(1), 7–42. doi
  6. Oota, S. R., Marreddy, M., Gupta, M., & Bapi, R. S. (2023). How does the brain process syntactic structure while listening? Findings of ACL 2023. arXiv
  7. Reddy, A. J., & Wehbe, L. (2021). Can fMRI reveal the representation of syntactic structure in the brain? NeurIPS 34. proceedings
  8. Oota, S. R., Gupta, M., & Toneva, M. (2023). Joint processing of linguistic properties in brains and language models. NeurIPS 36. arXiv
  9. He, L., Zhong, T., Antonello, R., Mischler, G., Goldblum, M., & Mesgarani, N. (2025). Far from the shallow: Brain-predictive reasoning embedding through residual disentanglement. NeurIPS 2025. arXiv
  10. Hadidi, N., Feghhi, E., Song, B. H., Blank, I. A., & Kao, J. C. (2026). Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds. Nature Communications, 17, 5769. doi
  11. Toneva, M., Mitchell, T. M., & Wehbe, L. (2022). Combining computational controls with natural text reveals aspects of meaning composition. Nature Computational Science, 2, 745–757. doi
  12. Merlin, G., & Toneva, M. (2024). Language models and brains align due to more than next-word prediction and word-level information. EMNLP 2024. arXiv
  13. Oota, S. R., Çelik, E., Deniz, F., & Toneva, M. (2024). Speech language models lack important brain-relevant semantics. ACL 2024. arXiv
  14. LeBel, A., Wagner, L., Jain, S., Adhikari-Desai, A., Gupta, B., Morgenthal, A., Tang, J., Xu, L., & Huth, A. G. (2023). A natural language fMRI dataset for voxelwise encoding models. Scientific Data, 10, 555. doi
  15. Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength natural language processing in Python. Zenodo. doi
  16. Benara, V., Singh, C., Morris, J. X., Antonello, R., Stoica, I., Huth, A. G., & Gao, J. (2024). Crafting interpretable embeddings by asking LLMs questions. NeurIPS 2024. arXiv
  17. Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., & Biderman, S. (2023). LEACE: Perfect linear concept erasure in closed form. NeurIPS 36. arXiv