الباحثون

Prasanna Sattigeri

المنشورات 2

نسخة أولية وصول مفتوح

Learning from Teacher Continuations at Student States

Haojin Wang, Dylan Zhang, Huaibo Chen وآخرون · 2026

We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillatio …

نسخة أولية وصول مفتوح

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge only after multiple turns, yet existing evaluations often reduce agent behavior to task or a …

المؤلفون المشاركون