Preprint Open access
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We o …