[
    {
        "id": "osp-16325",
        "type": "article-journal",
        "title": "VioLA: Learning Generalist Humanoid Control Policies from Human Data",
        "author": [
            {
                "family": "Albaba",
                "given": "Mert"
            },
            {
                "family": "Beißwenger",
                "given": "Jens"
            },
            {
                "family": "Manasyan",
                "given": "Anna"
            },
            {
                "family": "Marta",
                "given": "Daniel"
            },
            {
                "family": "Black",
                "given": "Michael J."
            },
            {
                "family": "Brendel",
                "given": "Wieland"
            },
            {
                "family": "Krause",
                "given": "Andreas"
            },
            {
                "family": "Martius",
                "given": "Georg"
            },
            {
                "family": "Riedmiller",
                "given": "Martin"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/viola-learning-generalist-humanoid-control-policies-from-human-data",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released."
    }
]