Abstract

We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level ("sensory") lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM's data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT's three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard's NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Brasoveanu, A., Takmaz, E., & Dotlačil, J. (2026). Relational Attention for Data-Efficient Language Modeling. https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling

MLA 9

Brasoveanu, Adrian, et al. "Relational Attention for Data-Efficient Language Modeling." https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling.

Chicago (author–date)

Brasoveanu, Adrian, Ece Takmaz, and Jakub Dotlačil. 2026. "Relational Attention for Data-Efficient Language Modeling." https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling.

Harvard

Brasoveanu, A., Takmaz, E. and Dotlačil, J. (2026) 'Relational Attention for Data-Efficient Language Modeling', Available at: https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling.

Vancouver

Brasoveanu A, Takmaz E, Dotlačil J. Relational Attention for Data-Efficient Language Modeling. https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling

IEEE

A. Brasoveanu, E. Takmaz, and J. Dotlačil, "Relational Attention for Data-Efficient Language Modeling," https://omanscience.com/en/articles/relational-attention-for-data-efficient-language-modeling.