الباحثون

Hai "Helen" Li

المنشورات 4

نسخة أولية وصول مفتوح

Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation

Lichen Zhu, Yiheng Wang, Yueqian Lin وآخرون · 2026

Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring …

نسخة أولية وصول مفتوح

Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

Lichen Zhu, Yueqian Lin, Yiheng Wang وآخرون · 2026

Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the fra …

نسخة أولية وصول مفتوح

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

Lichen Zhu, Yueqian Lin, Yiheng Wang وآخرون · 2026

Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop …

نسخة أولية وصول مفتوح

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

Zhixu Du, Weijia Han, Hai "Helen" Li وآخرون · 2026

Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually re …

المؤلفون المشاركون