Abstract
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lek, G., Xia, Z., Chen, P. Y., & Chen, L. Y. (2026). Understanding Confabulation and Rethinking Reconstruction in Activation Explanations. https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations
MLA 9
Lek, Gert, et al. "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations." https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
Chicago (author–date)
Lek, Gert, Zixuan Xia, Pin-Yu Chen, and Lydia Y. Chen. 2026. "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations." https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
Harvard
Lek, G., Xia, Z., Chen, P. Y. and Chen, L. Y. (2026) 'Understanding Confabulation and Rethinking Reconstruction in Activation Explanations', Available at: https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
Vancouver
Lek G, Xia Z, Chen PY, Chen LY. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations. https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations
IEEE
G. Lek, Z. Xia, P. Y. Chen, and L. Y. Chen, "Understanding Confabulation and Rethinking Reconstruction in Activation Explanations," https://omanscience.com/en/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.