الباحثون

Yang Li

المنشورات 15

نسخة أولية وصول مفتوح

VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile Manipulation

Yutian Zhang, Xingrui Xiong, Siyuan Ma وآخرون · 2026

Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing h …

نسخة أولية وصول مفتوح

Self-Evaluating Recursive Agents

TianYi Lyu, Xiaozhe Li, Yang Li وآخرون · 2026

Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and carry no ground truth. Existing methods score each node with …

نسخة أولية وصول مفتوح

SecJev: Bringing Security Expertise to System One Decision Models

Zheng Chen, Fei Yu, Haohao Huang وآخرون · 2026

Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge the first family of Je …

نسخة أولية وصول مفتوح

FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering

Haifeng Huang, Biyin Xu, Chunsheng Xin وآخرون · 2026

Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, uniformly sampled candidate pool, preventing evidence outside this pool from ever being selec …

نسخة أولية وصول مفتوح

Diagnosing On-Policy Self-Distillation for Reasoning Language Models

Yang Li, Gongle Xue, Yuheng Yuan وآخرون · 2026

On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavi …

نسخة أولية وصول مفتوح

All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping

Yang Li, Aming WU, Zihao Zhang وآخرون · 2026

To advance the development of embodied intelligence, Open-Environment Active 3D Mapping has attracted increasing attention, aiming to perform a long-horizon and shortest trajectory exploration for reconstructing unseen scenarios. Since only limited information about unseen environments is available, methods built on th …

نسخة أولية وصول مفتوح

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Shifeng Bao, Fanding Huang, Yihan Lin وآخرون · 2026

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha …

نسخة أولية وصول مفتوح

Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation

Yang Li, Sijia Zhang, Yihan Li وآخرون · 2026

Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predict …

نسخة أولية وصول مفتوح

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

Zhiqiang Xia, Yang Li, Xinyuan Zhang وآخرون · 2026

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i …

نسخة أولية وصول مفتوح

Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

Junlin Chen, Daize Dong, Huanwei Di وآخرون · 2026

BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient …

نسخة أولية وصول مفتوح

Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study

Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training …

نسخة أولية وصول مفتوح

Seeing and Solving Are Not Enough for Vision-Language Models

Ziheng Wang, Mingxuan Xie, Yilin Liu وآخرون · 2026

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separ …

نسخة أولية وصول مفتوح

HarnessPAI: An Evolving Harness for Physical AI

Xin Wang, Wenhao Wu, Menghao Zhang وآخرون · 2026

Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabi …

المؤلفون المشاركون