Preprint Open access
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape sub …
Preprint Open access
How can we personalize a shared expert library from a user's pairwise choices? Prior work can realize different reward trade-offs by merging reward-specialized experts, given a vector of trade-off weights. In practice, users can more naturally choose between outputs than specify numerical weights. The challenge is ther …
Preprint Open access
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that s …
Preprint Open access
Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while trainin …