الباحثون

Xin Chen

المنشورات 9

نسخة أولية وصول مفتوح

V-CoLA: Vision Token Compression with Linear Attention

Hao Jiang, Yiru Mao, Tianpeng Bu وآخرون · 2026

Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating …

نسخة أولية وصول مفتوح

CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training

Jing Li, Jian Meng, Yingmeng Gao وآخرون · 2026

Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the train …

نسخة أولية وصول مفتوح

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

Haifeng Wu, Srinivasan Manoharan, Jian Wan وآخرون · 2026

Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distill …

نسخة أولية وصول مفتوح

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Mei Wu, Rui Xie, Runyu Zhang وآخرون · 2026

Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-p …

نسخة أولية وصول مفتوح

Xiaomi-OCR-0 Technical Report

Xin Chen, Anan Du, Feng Feng وآخرون · 2026

Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-c …

نسخة أولية وصول مفتوح

Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment

Shunchang Liu, Lukas Fluri, Xin Chen وآخرون · 2026

Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, lea …

نسخة أولية وصول مفتوح

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Liyang Fan, Yingcheng Shi, Yongbin Li وآخرون · 2026

Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluations do not show whether those systems improve executable repository work. Repository benchmarks test code changes but do not isolate memory, while memory benchmarks score r …

نسخة أولية وصول مفتوح

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

Xin Chen, Sen Chen, Yujuan Ding وآخرون · 2026

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed …

المؤلفون المشاركون