Abstract

Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears. Detecting such activation is difficult because malicious behavior can comprise individually plausible actions, while unfamiliar tasks introduce legitimate changes in observations and behavior. We introduce TMT, a runtime backdoor detector based on Token Manifold and latent Transition modeling. Trained on benign rollouts, its two branches assess input-token structure and prediction errors in adjacent-layer latent dynamics. A suspicious rollout identified by the token manifold branch, once confirmed through latent deviations, guides transition selection for subsequent monitoring. We further explore policy purification through self-distillation: a frozen copy of the backdoored policy provides benign-input actions to supervise a student on paired benign and triggered observations, without requiring a separate clean reference policy. For evaluation, we adapt traditional backdoor detectors and repurpose anomaly and failure detection methods as VLA backdoor detectors. In a post-hoc comparison with ten baselines, TMT achieves state-of-the-art backdoor detection performance on unseen tasks across three VLA backdoor attacks. Our project page is available at https://zzr42.github.io/tmt/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhou, Z., Zhang, J., Xu, H., Zhang, X., Wen, E., Sun, J., & Jia, H. (2026). TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks. https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks

MLA 9

Zhou, Zirun, et al. "TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks." https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks.

Chicago (author–date)

Zhou, Zirun, Jingfeng Zhang, HaoChuan Xu, Xizhe Zhang, Elliott Wen, Jing Sun, and Hong Jia. 2026. "TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks." https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks.

Harvard

Zhou, Z., Zhang, J., Xu, H., Zhang, X., Wen, E., Sun, J. and Jia, H. (2026) 'TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks', Available at: https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks.

Vancouver

Zhou Z, Zhang J, Xu H, Zhang X, Wen E, Sun J, et al. TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks. https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks

IEEE

Z. Zhou, J. Zhang, H. Xu, X. Zhang, E. Wen, J. Sun, and H. Jia, "TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks," https://omanscience.com/en/articles/tmt-runtime-backdoor-detection-for-vision-language-action-policies-on-unseen-tasks.