الباحثون

Chris Russell

المنشورات 4

نسخة أولية وصول مفتوح

Settle: Learning When to Stop Reasoning

Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwe …

نسخة أولية وصول مفتوح

Decoupling Token Roles in Autoregressive Pretraining

Suqin Yuan, Runqi Lin, Kevin Qinghong Lin وآخرون · 2026

Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for …

نسخة أولية وصول مفتوح

SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

Ziwen Li, Hanlue Zhang, Zhenyang Ren وآخرون · 2026

Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic …

نسخة أولية وصول مفتوح

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

Suqin Yuan, Runqi Lin, Muyang Li وآخرون · 2026

Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show t …

المؤلفون المشاركون