الباحثون

Xinyu Zhang

المنشورات 11

نسخة أولية وصول مفتوح

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu وآخرون · 2026

Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade mode …

نسخة أولية وصول مفتوح

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

Shangye Song, Dong Gong, Hong Jia وآخرون · 2026

Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet exi …

نسخة أولية وصول مفتوح

Confidence-Ordering Reversal under Contextual Priors in Neural Decoding

Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. We study how a prior shapes confidence in speech retrieval on MEG-MASC an …

نسخة أولية وصول مفتوح

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Ziqun Bao, Xinyu Zhang, Yuchen Shao وآخرون · 2026

Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in check …

نسخة أولية وصول مفتوح

Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife

Yuzhuo Li, Di Zhao, Xinyu Zhang وآخرون · 2026

Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a singl …

نسخة أولية وصول مفتوح

End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

Shenglan Li, Rui Yao, Kunyang Sun وآخرون · 2026

RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-lab …

نسخة أولية وصول مفتوح

MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing

Zeyu Yang, Xinyu Zhang, Zibo Bi وآخرون · 2026

Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains. A balanced …

نسخة أولية وصول مفتوح

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

Xize Cheng, Wenxu Jia, Chenyuhao Wen وآخرون · 2026

Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to prese …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

المؤلفون المشاركون