الباحثون

Muhammad Ahmed Mohsin

المنشورات 3

نسخة أولية وصول مفتوح

Monte Carlo Estimation for KV Cache Eviction

Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fix …

نسخة أولية وصول مفتوح

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

Jintao Huang, Yifan Wang, Hongyu Shen وآخرون · 2026

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b …

نسخة أولية وصول مفتوح

Teaching Agents to Code Reliably

Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather …

المؤلفون المشاركون