Preprint Open access
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different id …