الباحثون

Yue Wang

المنشورات 20

نسخة أولية وصول مفتوح

PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies

Qing Huang, Yifei Yang, Ziqing Zou وآخرون · 2026

Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is t …

نسخة أولية وصول مفتوح

From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning

Fengyuan Liu, Yue Wang, Hangxi Guo وآخرون · 2026

Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids co …

نسخة أولية وصول مفتوح

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

Hangxi Guo, Fengyuan Liu, Yue Wang وآخرون · 2026

Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on fe …

نسخة أولية وصول مفتوح

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Hexuan Deng, Yue Wang, Wenyu Jiang وآخرون · 2026

Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repe …

نسخة أولية وصول مفتوح

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

Hexuan Deng, Zihao Yan, Xuebo Liu وآخرون · 2026

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a …

نسخة أولية وصول مفتوح

EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs

Leizhen Wang, Peibo Duan, Zhenlin Qin وآخرون · 2026

Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rul …

نسخة أولية وصول مفتوح

Learning Modular Policy for Multi-Floor Object Navigation:A Factorized Framework for Diagnostic Study

Shichao Zhai, Shuhao Ye, Rong Xiong وآخرون · 2026

Object-goal navigation (ObjectNav) in multi-floor scenarios presents a challenge due to sparse rewards caused by long-horizon decision-making. In this paper, we propose a diagnostic study based on a modular framework with an effective learnable policy to analyze failure factors in multi-floor scenarios. To achieve an e …

نسخة أولية وصول مفتوح

eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing

Dehao Huang, Jianbang Liu, Jianpan Gao وآخرون · 2026

Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of …

نسخة أولية وصول مفتوح

Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

Kemou Li, Qizhou Wang, Yue Wang وآخرون · 2026

Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against …

نسخة أولية وصول مفتوح

Robust Nash Alignment under Preference Uncertainty

Shihab Ahmed, Debamita Ghosh, David Tang وآخرون · 2026

Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain …

نسخة أولية وصول مفتوح

Collision-Aware and Observation-Aligned Object-Centric Scene Reconstruction from Point Cloud

Yuxuan Xie, Xuan Yu, Rong Xiong وآخرون · 2026

Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity and insufficient geometric constraints. We propose COOL, a fra …

نسخة أولية وصول مفتوح

BIND: Binding 3D Robot Actions to 2D Image Features

Cameron Smith, Arsh Tangri, Vitor Guizilini وآخرون · 2026

We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically …

نسخة أولية وصول مفتوح

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Hongyang Du, Yunfei Xie, Junjie Ye وآخرون · 2026

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choi …

نسخة أولية وصول مفتوح

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao وآخرون · 2026

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formula …

نسخة أولية وصول مفتوح

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

Weizhen Bian, Sitong Cheng, Rongxiu Zhong وآخرون · 2026

Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational con …

نسخة أولية وصول مفتوح

The Neverwhere Visual Parkour Benchmark Suite

Ziyu Chen, Henghui Bao, Haoran Chang وآخرون · 2026

State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop …

المؤلفون المشاركون