Abstract

Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lek, G., Xia, Z., Chen, P. Y., & Chen, L. Y. (2026). Understanding Confabulation and Rethinking Reconstruction in Activation Explanations. https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations

MLA 9

Lek, Gert, et al. "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations." https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

Chicago (author–date)

Lek, Gert, Zixuan Xia, Pin-Yu Chen, and Lydia Y. Chen. 2026. "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations." https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

Harvard

Lek, G., Xia, Z., Chen, P. Y. and Chen, L. Y. (2026) 'Understanding Confabulation and Rethinking Reconstruction in Activation Explanations', Available at: https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

Vancouver

Lek G, Xia Z, Chen PY, Chen LY. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations. https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations

IEEE

G. Lek, Z. Xia, P. Y. Chen, and L. Y. Chen, "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations," https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.