Abstract

Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hu, K., Hu, S., Hong, F., Chen, Z., & Liu, Z. (2026). Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement. https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement

MLA 9

Hu, Kairui, et al. "Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement." https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement.

Chicago (author–date)

Hu, Kairui, Siyuan Hu, Fangzhou Hong, Zhaoxi Chen, and Ziwei Liu. 2026. "Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement." https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement.

Harvard

Hu, K., Hu, S., Hong, F., Chen, Z. and Liu, Z. (2026) 'Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement', Available at: https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement.

Vancouver

Hu K, Hu S, Hong F, Chen Z, Liu Z. Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement. https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement

IEEE

K. Hu, S. Hu, F. Hong, Z. Chen, and Z. Liu, "Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement," https://omanscience.com/en/articles/embodied-turing-machines-stateful-code-for-robot-recursive-self-improvement.