Abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Torrielli, F., Barmina, G., Núñez, A. B., Rapp, A., Caro, L. D., Schneider-Kamp, P., & Poech, L. G. (2026). Selecting The Most Informative Tokens in Natural Language Autoencoders. https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders
MLA 9
Torrielli, Federico, et al. "Selecting The Most Informative Tokens in Natural Language Autoencoders." https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders.
Chicago (author–date)
Torrielli, Federico, Gianluca Barmina, Andrea Blasi Núñez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, and Lukas Galke Poech. 2026. "Selecting The Most Informative Tokens in Natural Language Autoencoders." https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders.
Harvard
Torrielli, F., Barmina, G., Núñez, A. B., Rapp, A., Caro, L. D., Schneider-Kamp, P. and Poech, L. G. (2026) 'Selecting The Most Informative Tokens in Natural Language Autoencoders', Available at: https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders.
Vancouver
Torrielli F, Barmina G, Núñez AB, Rapp A, Caro LD, Schneider-Kamp P, et al. Selecting The Most Informative Tokens in Natural Language Autoencoders. https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders
IEEE
F. Torrielli, G. Barmina, A. B. Núñez, A. Rapp, L. D. Caro, P. Schneider-Kamp, and L. G. Poech, "Selecting The Most Informative Tokens in Natural Language Autoencoders," https://omanscience.com/en/articles/selecting-the-most-informative-tokens-in-natural-language-autoencoders.