الباحثون

Wei Yang

المنشورات 12

نسخة أولية وصول مفتوح

Characterizing Overconfident Failure in LLM-Based Code Generation

Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or app …

نسخة أولية وصول مفتوح

SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning

Wei Yang, Shawn Li, Yuehan Qin وآخرون · 2026

Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipeli …

نسخة أولية وصول مفتوح

Backward-State Policy Is Part of the Learning Algorithm

Shuxiao Xie, Shuyang Xie, Dezhi Ran وآخرون · 2026

Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is …

نسخة أولية وصول مفتوح

Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication

Shuxiao Xie, Shuyang Xie, Yuan Cao وآخرون · 2026

Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on it …

نسخة أولية وصول مفتوح

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Zixuan Yang, Yiqun Chen, Qi Liu وآخرون · 2026

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff …

نسخة أولية وصول مفتوح

Controlled Decoding Attacks on Black-Box LLMs

Jesson Wang, Shawn Li, Wei Yang وآخرون · 2026

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampl …

نسخة أولية وصول مفتوح

From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction

Zixiao Dong, Wei Yang, Zihao Liu وآخرون · 2026

Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study …

نسخة أولية وصول مفتوح

Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction

Zixiao Dong, Wei Yang, Zihao Liu وآخرون · 2026

A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an …

نسخة أولية وصول مفتوح

How Well Do Pseudo-Rigid Body Models Capture Real Plants? In-Field Validation of Simulated Blueberry Canes

Hannah Kolano, Chelse VanAtter, Wei Yang وآخرون · 2026

Many labor-intensive tasks in fruit production such as pruning require physically interacting with the plant (e.g., pushing, pulling, bending limbs, etc.). Due to increasing labor shortages, there is widespread interest in the adoption of robotics in this area. When the robot must physically interact with the system, a …

المؤلفون المشاركون