Abstract

Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Pierucci, F., Syrnikov, M. B., Prandi, M., Galisai, M., Giarrusso, F., & Bisconti, P. (2026). Xeno-Interpretability: Investigating the Alien Minds of LLMs. https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms

MLA 9

Pierucci, F., et al. "Xeno-Interpretability: Investigating the Alien Minds of LLMs." https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms.

Chicago (author–date)

Pierucci, F., M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, and P. Bisconti. 2026. "Xeno-Interpretability: Investigating the Alien Minds of LLMs." https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms.

Harvard

Pierucci, F., Syrnikov, M. B., Prandi, M., Galisai, M., Giarrusso, F. and Bisconti, P. (2026) 'Xeno-Interpretability: Investigating the Alien Minds of LLMs', Available at: https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms.

Vancouver

Pierucci F, Syrnikov MB, Prandi M, Galisai M, Giarrusso F, Bisconti P. Xeno-Interpretability: Investigating the Alien Minds of LLMs. https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms

IEEE

F. Pierucci, M. B. Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, and P. Bisconti, "Xeno-Interpretability: Investigating the Alien Minds of LLMs," https://omanscience.com/en/articles/xeno-interpretability-investigating-the-alien-minds-of-llms.