الباحثون

Haitham Bou-Ammar

المنشورات 7

نسخة أولية وصول مفتوح

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

Haoyu Zhao, Zhengxu Yu, Zhiyuan He وآخرون · 2026

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on th …

نسخة أولية وصول مفتوح

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu وآخرون · 2026

A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves o …

نسخة أولية وصول مفتوح

Tail-Influence Sampling for CVaR Policy Evaluation

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es …

نسخة أولية وصول مفتوح

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

Matthieu Zimmer, Xiaotong Ji, Tu Nguyen وآخرون · 2026

Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answe …

نسخة أولية وصول مفتوح

Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents

Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action lon …

نسخة أولية وصول مفتوح

Composable Decoding on the Probability Simplex: Theory and Implementation

Decoding for large language models is typically treated as a collection of isolated sampling strategies, with limited theoretical understanding of the behaviours they induce and how their underlying objectives relate. We formulate decoding as an optimisation problem over next-token distributions on the probability simp …

المؤلفون المشاركون