الباحثون

Michal Valko

المنشورات 2

نسخة أولية وصول مفتوح

Monte Carlo Estimation for KV Cache Eviction

Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fix …

نسخة أولية وصول مفتوح

An analysis of Mirror-Descent Soft Actor-Critic

Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare …

المؤلفون المشاركون