[
    {
        "id": "osp-19407",
        "type": "article-journal",
        "title": "Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking",
        "author": [
            {
                "family": "Kocielnik",
                "given": "Rafal"
            },
            {
                "family": "Song",
                "given": "Peiyang"
            },
            {
                "family": "Han",
                "given": "Pengrui"
            },
            {
                "family": "Marmarelis",
                "given": "Myrl G."
            },
            {
                "family": "Debnath",
                "given": "Ramit"
            },
            {
                "family": "Mobbs",
                "given": "Dean"
            },
            {
                "family": "Alvarez",
                "given": "R. Michael"
            }
        ],
        "URL": "https://omanscience.com/en/articles/representational-control-over-self-report-behavior-coherence-in-llm-risk-taking",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation."
    }
]