[
    {
        "id": "osp-19564",
        "type": "article-journal",
        "title": "AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning",
        "author": [
            {
                "family": "Yu",
                "given": "Han"
            },
            {
                "family": "Zhu",
                "given": "Wenhui"
            },
            {
                "family": "Chen",
                "given": "Xiwen"
            },
            {
                "family": "Wang",
                "given": "Zhipeng"
            },
            {
                "family": "Sang",
                "given": "Hejian"
            },
            {
                "family": "Shi",
                "given": "Han"
            },
            {
                "family": "Zhou",
                "given": "Menglin"
            },
            {
                "family": "Dong",
                "given": "Xuanzhao"
            },
            {
                "family": "Huang",
                "given": "Minzhou"
            },
            {
                "family": "Cai",
                "given": "Rui"
            },
            {
                "family": "Wang",
                "given": "Hao"
            },
            {
                "family": "Geramifard",
                "given": "Alborz"
            }
        ],
        "URL": "https://omanscience.com/en/articles/avokv-e-payload-aware-kv-cache-eviction-for-long-reasoning",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory."
    }
]