Preprint Open access
Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models
Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to prese …