Authors

Adel Bibi

Publications 3

Preprint Open access

Towards a Unified Misuse Monitoring Benchmark

LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats s …

Preprint Open access

Towards Better Exploration in Sequential Test-Time Scaling

Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to sol …

Co-authors