الباحثون

Tu Nguyen

المنشورات 2

نسخة أولية وصول مفتوح

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu وآخرون · 2026

A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves o …

نسخة أولية وصول مفتوح

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

Matthieu Zimmer, Xiaotong Ji, Tu Nguyen وآخرون · 2026

Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answe …

المؤلفون المشاركون