الباحثون

Sahar Abdelnabi

المنشورات 4

نسخة أولية وصول مفتوح

Surviving the Router: Optimizing Skill Injections for Retrieval and Execution

Haneen Najjar, Luca Scionis, Haritz Puerto وآخرون · 2026

AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execu …

نسخة أولية وصول مفتوح

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into la …

نسخة أولية وصول مفتوح

Training LLMs to Verbalize Evaluation Awareness

Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supe …

المؤلفون المشاركون