الباحثون

Long Chen

المنشورات 11

نسخة أولية وصول مفتوح

Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks

Zhewei Fang, Yuxin Zhang, Zhenwei Shao وآخرون · 2026

Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs …

نسخة أولية وصول مفتوح

Human Behavior-Informed Crash Scenario Generation with Real-World Crash Priors for Autonomous Vehicle Safety Evaluation

Mingxing Peng, Xusen Guo, Long Chen وآخرون · 2026

Reliable safety evaluation of autonomous vehicles (AVs) is essential to improving road safety, yet it depends critically on realistic simulation of rare crashes. Existing crash scenario generation methods can increase collision occurrence, but often fail to realistically reproduce how crashes evolve before impact or th …

نسخة أولية وصول مفتوح

Smoother Flow Matching via Contrastive Trajectory Repulsion

Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the interm …

نسخة أولية وصول مفتوح

RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

Yuzhou Wu, Longteng Fan, Zimeng Li وآخرون · 2026

Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and comple …

نسخة أولية وصول مفتوح

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Guannan Lai, Gelin Bian, Hao-Xuan Ma وآخرون · 2026

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision …

نسخة أولية وصول مفتوح

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

Bangjun Wang, Longyan Wu, Yukun Wei وآخرون · 2026

Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-a …

نسخة أولية وصول مفتوح

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Wei Chen, Xuanyu Zheng, Yancheng Long وآخرون · 2026

Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and t …

نسخة أولية وصول مفتوح

ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision

Yujian Yuan, Xin Cai, Yufan Chen وآخرون · 2026

Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection and correction before …

نسخة أولية وصول مفتوح

Causal-History Test-Time Scaling for Failure Recovery in Autoregressive World-Action Models

Lin Li, Long Chen, Kwunhang وآخرون · 2026

World-action models (WAMs) have emerged as a promising paradigm for robot manipulation by jointly modeling future visual dynamics and robot actions. However, existing WAMs are trained predominantly on successful trajectories, making them prone to failure when real-world execution diverges from the learned dynamics. Thi …

المؤلفون المشاركون