Authors

Yang Li

Publications 15

Preprint Open access

Self-Evaluating Recursive Agents

TianYi Lyu, Xiaozhe Li, Yang Li et al. · 2026

Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and carry no ground truth. Existing methods score each node with …

Preprint Open access

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i …

Preprint Open access

Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study

Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training …

Preprint Open access

Seeing and Solving Are Not Enough for Vision-Language Models

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separ …

Preprint Open access

HarnessPAI: An Evolving Harness for Physical AI

Xin Wang, Wenhao Wu, Menghao Zhang et al. · 2026

Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabi …

Co-authors