Abstract
Linear probing and activation steering use linear directions to predict and control behavior-level concepts such as correctness, safety, and social bias in question answering. We call these \emph{behavior-level concept directions}. Drawing on the Linear Representation Hypothesis (LRH) and circuit studies, these directions are often interpreted as concept representations. Yet controlled evidence for linear concept structure and local mechanisms does not establish that these directions represent the intended behavioral concepts, leaving probing and steering without a unified theoretical account. We propose the \emph{Answer-Basin Representation Hypothesis} (ABRH): the model's own answer measure organizes the linear structure of these directions. For each question, all continuations yielding the same answer form an answer basin, whose mass is their total probability; these masses define the model's answer distribution. ABRH posits that its concentration before generation and the relative mass of each answer after generation are represented along linear directions shared across questions. Experiments span four Qwen2.5 and Gemma-3 models on correctness, social bias, and safety tasks. Decoupling concept labels from basin mass shows that probing and steering exhibit concept-consistent effects when labels align with mass orderings, weaken as mass gaps shrink, and reverse under conflict. Probes trained solely on mass orderings among wrong answers still select correct answers; a fixed steering vector can instead favor wrong answers when labels conflict with mass orderings. These results support an answer-measure account of behavior-level concept directions and of when probing and steering succeed, fail, or reverse.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yu, M., Li, H., Wang, Z., Chen, J., Li, X., Singh, P., Cao, Y., & Hu, L. (2026). The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts. https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts
MLA 9
Yu, Manjiang, et al. "The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts." https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts.
Chicago (author–date)
Yu, Manjiang, Hongji Li, Zihan Wang, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, and Lijie Hu. 2026. "The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts." https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts.
Harvard
Yu, M., Li, H., Wang, Z., Chen, J., Li, X., Singh, P., Cao, Y. and Hu, L. (2026) 'The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts', Available at: https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts.
Vancouver
Yu M, Li H, Wang Z, Chen J, Li X, Singh P, et al. The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts. https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts
IEEE
M. Yu, H. Li, Z. Wang, J. Chen, X. Li, P. Singh, Y. Cao, and L. Hu, "The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts," https://omanscience.com/en/articles/the-answer-basin-representation-hypothesis-we-are-not-probing-or-steering-concepts.