الباحثون

Marco Pavone

المنشورات 18

نسخة أولية وصول مفتوح

Parametric Trajectory Distillation for Few-Step Video Generation

Lan Feng, Peter Karkus, Maximilian Igl وآخرون · 2026

Video diffusion and flow models require many sequential evaluations, making generation computationally expensive. Few-step distillation reduces this cost but poses a capacity allocation problem: a student must match the teacher's iterative generation with far less sequential computation. Existing trajectory methods ask …

نسخة أولية وصول مفتوح

Co-Evolving Robot Orchestrators and Policies through Deployment

Xilun Zhang, Maggie Wang, Erik Bauer وآخرون · 2026

Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when t …

نسخة أولية وصول مفتوح

Grounding What Shapes the Plan: Rethinking Groundedness for Physical Intelligence in Autonomous Driving

Minkyoung Cho, Zewei Zhou, Wenhao Ding وآخرون · 2026

Driving models increasingly ground reasoning in causal relations, spatial structure, perceptual evidence, and predicted futures. These advances make reasoning more faithful to the driving scene, but leave a fundamental question unresolved: what should groundedness mean when the model ultimately outputs an action? Corre …

نسخة أولية وصول مفتوح

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

Chensheng Peng, Wenhao Ding, Ran Tian وآخرون · 2026

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT reco …

نسخة أولية وصول مفتوح

OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

Yuji Takubo, Daniele Gammelli, Marco Pavone وآخرون · 2026

Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an …

نسخة أولية وصول مفتوح

DeepJEPA: Scaling World Models from Within

Zijian Jin, Yunbei Zhang, Yuanzhe Liu وآخرون · 2026

World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrate …

نسخة أولية وصول مفتوح

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

Yunbei Zhang, Zijian Jin, Yuanzhe Liu وآخرون · 2026

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu …

نسخة أولية وصول مفتوح

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk وآخرون · 2026

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show …

نسخة أولية وصول مفتوح

A Scene Language Model for Open-Vocabulary Scene Mapping

Adam Lilja, Fabio Hübel, Siming He وآخرون · 2026

Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store featur …

نسخة أولية وصول مفتوح

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

Damiano Da Col, Maximilian Igl, Peter Karkus وآخرون · 2026

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop de …

نسخة أولية وصول مفتوح

ElastiQP: An Always-Feasible QP Solver for Constrained Robot Control

As robot capabilities increase, quadratic programming (QP)-based controllers must account for a similarly increasing number of constraints to ensure safe, reliable operation. Yet, with each added constraint, this introduces more chances of momentary conflict: in which case, a QP solver that returns an "infeasible" stat …

المؤلفون المشاركون