Authors

Quansheng Gu

Publications 1

Preprint Open access

SparseEngine: Sparse-First Inference Engine

Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specif …

Co-authors