الباحثون

Saiyue Lyu

المنشورات 2

نسخة أولية وصول مفتوح

Safety Must Survive Self-Improvement: Why Failures Persist and How Agents Recover

Yunbei Zhang, Janet Wang, Saiyue Lyu وآخرون · 2026

Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization task …

نسخة أولية وصول مفتوح

Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems

Yunbei Zhang, Saiyue Lyu, Janet Wang وآخرون · 2026

Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive action avoids disclosure but defeats the …

المؤلفون المشاركون