الباحثون

Xiaoying Tang

المنشورات 6

نسخة أولية وصول مفتوح

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

Mingda Zhang, Wenjin Liu, Tiesunlong Shen وآخرون · 2026

Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self- …

نسخة أولية وصول مفتوح

How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation

Haoyue Liu, Zhichao Wang, Huanyu Yan وآخرون · 2026

Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under ti …

نسخة أولية وصول مفتوح

How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation

Xiaoyu Ma, Haoyue Liu, Zhichao Wang وآخرون · 2026

On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effectively? Prior work shows that when student prefixes follow …

نسخة أولية وصول مفتوح

JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion

Zhengkai Tu, Mingda Zhang, Zijia Wang وآخرون · 2026

A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and ungrounded discretion, benchmarks score a label or a rubric, and the experience that would supply the balance stays unverified. We formalize legal ju …

نسخة أولية وصول مفتوح

EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment

Mingda Zhang, Hanwen Zhang, Qiang Huang وآخرون · 2026

In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion tha …

نسخة أولية وصول مفتوح

CollabFlow: Recursive Self-Improvement of Agent Collaboration

Xiao Huang, Mingda Zhang, Junming Zhang وآخرون · 2026

Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator …

المؤلفون المشاركون