Preprint Open access
Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational ove …
Preprint Open access
Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet thes …
Preprint Open access
Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of e …