الباحثون

Haoyu Wang

المنشورات 17

نسخة أولية وصول مفتوح

Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery

Li Lu, Yanjie Zhao, Hongjie Chen وآخرون · 2026

LLM agents can spend millions of tokens during vulnerability discovery without producing a working proof of concept (PoC). What consumes that budget, and why does it fail to produce results? We diagnose these costs and failures through a multi-axis open-coding study of 200 CyberGym traces, spanning four agents (i.e., C …

نسخة أولية وصول مفتوح

RFPO: Rectified Flow Policy Optimization for Embodied Control

Ting Huang, Lisiyu Pan, Haoyu Wang وآخرون · 2026

Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain relia …

نسخة أولية وصول مفتوح

InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

Yuchen Li, Shaoyang Zhou, Yiran Wang وآخرون · 2026

Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separa …

نسخة أولية وصول مفتوح

Directed Temporal Representations for Offline Visual Control

Chenyang Yuan, Haoyu Wang, Zhuo Sun وآخرون · 2026

Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on t …

نسخة أولية وصول مفتوح

ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks

Qi Cheng, Rongchao Dong, Shengyu Chen وآخرون · 2026

The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark des …

نسخة أولية وصول مفتوح

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

Qi Cheng, Shengyu Chen, Wei Cheng وآخرون · 2026

Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly …

نسخة أولية وصول مفتوح

RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agentic Workflow Generation

Qi Cheng, Shengyu Chen, Wei Cheng وآخرون · 2026

Agentic workflows enable large language models (LLMs) to solve complex tasks by coordinating reasoning, tool use, and verification. However, a workflow optimized for an entire task collection can overlook differences in the reasoning strategies that individual queries need, while searching for a new workflow for every …

نسخة أولية وصول مفتوح

WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration

Qi Cheng, Shengyu Chen, Wei Cheng وآخرون · 2026

Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successful executions. We present WorkflowOps, a multi-agent workflow orchestration framework th …

نسخة أولية وصول مفتوح

Causal Improvement Graph for Agentic Harness Optimization

Junjie Zhang, Shunyu Liu, Haoyu Wang وآخرون · 2026

Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Exist …

نسخة أولية وصول مفتوح

Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

Xuehang Guo, Haoyu Wang, Shengyu Chen وآخرون · 2026

Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all …

نسخة أولية وصول مفتوح

It Takes Workflows to Evolve Better Workflows

Xuehang Guo, Haoyu Wang, Haifeng Chen وآخرون · 2026

Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only th …

نسخة أولية وصول مفتوح

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

Zheng Fang, Yongmin Li, Yichang Zhang وآخرون · 2026

Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behavior can become incre …

نسخة أولية وصول مفتوح

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

Haoyu Wang, Wei Zhao, Yedi Zhang وآخرون · 2026

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa …

نسخة أولية وصول مفتوح

SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

Ziwen Li, Hanlue Zhang, Zhenyang Ren وآخرون · 2026

Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic …

نسخة أولية وصول مفتوح

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Lujia Bao, Qian Chen, Luyao Cheng وآخرون · 2026

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis …

المؤلفون المشاركون