الباحثون

Shuo Yang

المنشورات 11

نسخة أولية وصول مفتوح

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Shuo Yang, Changbai Li, Linlin Yang وآخرون · 2026

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on ou …

نسخة أولية وصول مفتوح

ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs

Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after sel …

نسخة أولية وصول مفتوح

APEX: Active Protection at Execution Boundaries for LLM Agents

Xinran Zheng, Xin Fan Guo, Zhiqiang Hao وآخرون · 2026

Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind th …

نسخة أولية وصول مفتوح

Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

Shuo Yang, Lihao Fang, Yi Zhang وآخرون · 2026

Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tens …

نسخة أولية وصول مفتوح

EffGS: Efficient and High-Fidelity Gaussian Splatting

Changbai Li, Shuo Yang, Yichen Yang وآخرون · 2026

3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality degradation when extended to more complex, large-scale scenes. To address this issue, we propose EffGS, a more general acceleration framework that improves training and …

نسخة أولية وصول مفتوح

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Yizhao Li, Pusen Gao, Ming Wang وآخرون · 2026

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-alig …

نسخة أولية وصول مفتوح

GleanVID: Complementary Token Selection for Efficient Video Large Language Models

Shuo Yang, Changbai Li, Rui Tang وآخرون · 2026

Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and conseque …

نسخة أولية وصول مفتوح

Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

Xiangcheng Zhan, Zirui Chen, Yicheng Zhao وآخرون · 2026

World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimension …

نسخة أولية وصول مفتوح

Whole-Body UMI: Transferring UMI Manipulation Skills to Humanoid Whole-Body Manipulation via Real-Time Motion Generation

Yuxuan Nai, Leixin Chang, Liangjing Yang وآخرون · 2026

Collecting whole-body demonstrations for humanoid manipulation mostly relies on teleoperation, which is costly and hard to scale up. The Universal Manipulation Interface (UMI) provides a scalable data collection paradigm, but end-effector trajectories alone underdetermine humanoid whole-body coordination, which is insu …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

المؤلفون المشاركون