Abstract

Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Kocielnik, R., Song, P., Han, P., Marmarelis, M. G., Debnath, R., Mobbs, D., & Alvarez, R. M. (2026). Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking. https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking

MLA 9

Kocielnik, Rafal, et al. "Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking." https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking.

Chicago (author–date)

Kocielnik, Rafal, Peiyang Song, Pengrui Han, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, and R. Michael Alvarez. 2026. "Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking." https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking.

Harvard

Kocielnik, R., Song, P., Han, P., Marmarelis, M. G., Debnath, R., Mobbs, D. and Alvarez, R. M. (2026) 'Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking', Available at: https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking.

Vancouver

Kocielnik R, Song P, Han P, Marmarelis MG, Debnath R, Mobbs D, et al. Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking. https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking

IEEE

R. Kocielnik, P. Song, P. Han, M. G. Marmarelis, R. Debnath, D. Mobbs, and R. M. Alvarez, "Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking," https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking.