الباحثون

Yiran Chen

المنشورات 9

نسخة أولية وصول مفتوح

Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation

Lichen Zhu, Yiheng Wang, Yueqian Lin وآخرون · 2026

Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring …

نسخة أولية وصول مفتوح

Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

Lichen Zhu, Yueqian Lin, Yiheng Wang وآخرون · 2026

Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the fra …

نسخة أولية وصول مفتوح

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

Lichen Zhu, Yueqian Lin, Yiheng Wang وآخرون · 2026

Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop …

نسخة أولية وصول مفتوح

MARS: Multi-resolution Adaptive Routing for Sequential Recommendation

Ming Yin, Sixun Dong, Yudong Liu وآخرون · 2026

Long-history recommenders often compress each user's history into a compact, candidate-independent memory that is cached and reused to score large candidate pools. We show that real user histories exhibit multi-scale semantic structure, with short-lived intent, medium-term interests, and long-term preferences coexistin …

نسخة أولية وصول مفتوح

The Functional Structure of Post-Compression Recovery in Low-Rank LLMs

Zishan Shao, Liang Tian, Georgiy Zemlevskiy وآخرون · 2026

Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observed between methods at the compression endpoint may shrink, grow, or even reverse after recovery. We ask whether this recovery heterogeneity reflects …

نسخة أولية وصول مفتوح

Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation

Diffusion language model (dLLM) compression faces a known challenge because calibration is typically performed on clean, fully visible activations, whereas inference traverses partially masked intermediate states. For low-rank compression, this raises two questions. First, can low-rank optimality still be characterized …

نسخة أولية وصول مفتوح

Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance

Mingrun Jiang, Yuejia Liu, Zishan Shao وآخرون · 2026

Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activati …

نسخة أولية وصول مفتوح

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

Zhixu Du, Weijia Han, Hai "Helen" Li وآخرون · 2026

Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually re …

نسخة أولية وصول مفتوح

PQR3D: Progressive Query Refinement over Reference-Conditioned Temporal Windows for Multi-View 3D Object Detection

Hui Ye, Yudong Liu, Yiran Chen وآخرون · 2026

Temporal context is essential for camera-only multi-view 3D object detection. Existing streaming detectors maintain and propagate query states from one frame to the next, requiring sequence-aware training and chronological inference. We propose PQR3D, which performs progressive query refinement within referenceconditio …

المؤلفون المشاركون