الباحثون

Xin Li

المنشورات 17

نسخة أولية وصول مفتوح

VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

Runquan Gui, Hanzhu Chen, Zehao Wang وآخرون · 2026

Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence …

نسخة أولية وصول مفتوح

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

Wenqi Li, Mindi Ruan, Chuanbo Hu وآخرون · 2026

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~ …

نسخة أولية وصول مفتوح

MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code

Xueyi Chen, Shiyu Liu, Xin Jin وآخرون · 2026

Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce MetaKernelBench, which measures whether experience distilled f …

نسخة أولية وصول مفتوح

Loopy: Low-Bit Quantization Framework for Looped Language Models

Zeyu LI, Yipu ZHANG, Jintao Chen وآخرون · 2026

Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, but errors introduced by a quantized shared core affect subseq …

نسخة أولية وصول مفتوح

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Zhengming Yu, Junkun Yuan, Haotian Yang وآخرون · 2026

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adve …

نسخة أولية وصول مفتوح

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

Xinge Peng, Yiting Lu, Tianwu Zhi وآخرون · 2026

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inh …

نسخة أولية وصول مفتوح

S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

Jingdong Zhang, Xin Li, Jan Kautz وآخرون · 2026

Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forw …

نسخة أولية وصول مفتوح

ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing

Xijun Wang, Xin Li, Suhang Yao وآخرون · 2026

Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal cont …

نسخة أولية وصول مفتوح

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

Xijun Wang, Xin Li, Zirui Lang وآخرون · 2026

Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates referenc …

نسخة أولية وصول مفتوح

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We intro …

نسخة أولية وصول مفتوح

AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture

Shengjie Guo, Xin Li, Borjana Arsova وآخرون · 2026

Accurate counting and localization of plants and their organs support phenotyping and yield estimation, yet target appearance, scale, and density vary widely across species and imaging conditions. Exemplar boxes specify the target without category-specific retraining, and point predictions identify the individual insta …

نسخة أولية وصول مفتوح

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

Xin Li, Zhimin Mao, Shang Wang وآخرون · 2026

Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-guided, training-free aggregation method requiring neither place labels nor task-specific weight updates. Our key observ …

نسخة أولية وصول مفتوح

ForeTac-VLA: A Forecasting-Based Tactile-Vision-Language-Action Model for Contact-Rich Robotic Manipulation

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet their reliance on visual perception limits robustness in contact-rich environments, where critical physical interaction states may not be visually observable. Existing tactile-enhanced VLA methods improve physical gro …

نسخة أولية وصول مفتوح

RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents

Mengyi Deng, Xin Li, Duyi Pan وآخرون · 2026

Language agents increasingly rely on reusable skills, but post-failure repair is often handled by opaque one-shot reflection: a model generates a skill patch without explicitly maintaining how failure explanations relate to candidate repairs or how unsuccessful retests should influence later edits. We introduce RESKILL …

المؤلفون المشاركون