الباحثون

Yongliang Shen

المنشورات 8

نسخة أولية وصول مفتوح

SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

Hongxing Li, Jinyue Su, Dingming Li وآخرون · 2026

Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. …

نسخة أولية وصول مفتوح

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

Hongxing Li, Dingming Li, Yixin Li وآخرون · 2026

Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating v …

نسخة أولية وصول مفتوح

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Hongxing Li, Yixin Li, Dingming Li وآخرون · 2026

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spat …

نسخة أولية وصول مفتوح

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Yong Du, Tongbo Chen, Zhengxi Lu وآخرون · 2026

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privil …

نسخة أولية وصول مفتوح

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Xingyu Wu, Yuchen Yan, Zhengxi Lu وآخرون · 2026

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories int …

نسخة أولية وصول مفتوح

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

Bohao Wang, Chenwei Wu, Hang Zou وآخرون · 2026

Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However, existing telecom LLMs struggle to reliably reason across these diverse tasks and data types. Gen …

نسخة أولية وصول مفتوح

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Yan Yu, Zhengxi Lu, Yizhou Liu وآخرون · 2026

Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is und …

المؤلفون المشاركون