Preprint Open access
Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to …