الباحثون

Zhentao Tan

المنشورات 2

نسخة أولية وصول مفتوح

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Zhentao Tan, Jingyi Shen, Yanbo Li وآخرون · 2026

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogene …

نسخة أولية وصول مفتوح

SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models

Tianxiang Chen, Zhentao Tan, Zi Ye وآخرون · 2026

Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using …

المؤلفون المشاركون