Abstract

Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, T., Wang, H., Wan, J., Hu, T., & Wang, C. (2026). How the Audit Rule Shapes Faithful Factor Explanations in LLMs. https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms

MLA 9

Zhang, Taolin, et al. "How the Audit Rule Shapes Faithful Factor Explanations in LLMs." https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms.

Chicago (author–date)

Zhang, Taolin, Hanyu Wang, Jiuheng Wan, Tingyuan Hu, and Chengyu Wang. 2026. "How the Audit Rule Shapes Faithful Factor Explanations in LLMs." https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms.

Harvard

Zhang, T., Wang, H., Wan, J., Hu, T. and Wang, C. (2026) 'How the Audit Rule Shapes Faithful Factor Explanations in LLMs', Available at: https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms.

Vancouver

Zhang T, Wang H, Wan J, Hu T, Wang C. How the Audit Rule Shapes Faithful Factor Explanations in LLMs. https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms

IEEE

T. Zhang, H. Wang, J. Wan, T. Hu, and C. Wang, "How the Audit Rule Shapes Faithful Factor Explanations in LLMs," https://omanscience.com/en/articles/how-the-audit-rule-shapes-faithful-factor-explanations-in-llms.