الباحثون

Bowen Zhang

المنشورات 4

نسخة أولية وصول مفتوح

Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning

Zhijun Zhang, Qianlong Wang, Keyang Ding وآخرون · 2026

Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we r …

نسخة أولية وصول مفتوح

A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection

Bowen Zhang, Changrui Fang, Xinsong Ma وآخرون · 2026

Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the …

نسخة أولية وصول مفتوح

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

Maoqi Liu, Junwei He, Bowen Zhang وآخرون · 2026

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one anot …

نسخة أولية وصول مفتوح

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Bowen Zhang, Junwei He, Maoqi Liu وآخرون · 2026

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto …

المؤلفون المشاركون