الباحثون

Xin Wang

المنشورات 28

نسخة أولية وصول مفتوح

Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence

Yuchen Yang, Xin Wang, Lufan Wang وآخرون · 2026

Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiabl …

نسخة أولية وصول مفتوح

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge وآخرون · 2026

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac …

نسخة أولية وصول مفتوح

SkillPoison: Progressive Skill Poisoning via Successful Experiences

Lizhi Zhang, Xin He, Dianxuan Fu وآخرون · 2026

Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and …

نسخة أولية وصول مفتوح

Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering

Yang Hong, Yajun Yang, Xin Wang وآخرون · 2026

Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge sou …

نسخة أولية وصول مفتوح

Adaptive Power Sampling for LLM Reasoning

Bingnan Xiao, Chenhao Yang, Bingcong Li وآخرون · 2026

Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query diffic …

نسخة أولية وصول مفتوح

Collective Quantum State Preparation Is Nonadditive

Lin Zhu, Ranyiliu Chen, Xin Wang وآخرون · 2026

Quantum state preparation is a basic primitive of quantum computation and control, but its resource cost is usually analyzed for one target state or under an independent-copy construction. Here we study the many-copy problem under weighted Pauli control, where a Pauli-product rotation $e^{iφP}$ is charged by its rotati …

نسخة أولية وصول مفتوح

Closed-Form Arbitrary-Gap Rényi Nonadditivity Above Any Positive Order via Direct Products of Free Groups and High-Girth Graphs

We give a closed-form finite-dimensional construction of quantum channels with real algebraic entries exhibiting arbitrarily large violations of minimum-output Rényi entropy additivity uniformly above any prescribed positive Rényi order, using a different finite realization based on direct products of free groups and f …

نسخة أولية وصول مفتوح

EVISKILL: Grounding Skill Evolution in Replayable Evidence

Yan Zhou, Yili Wang, Yiwei Dai وآخرون · 2026

Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justified and when it should become persistent guidance. However, …

نسخة أولية وصول مفتوح

Towards Automatically Pruning Logging Code with Coding Agents: How Far Are We?

He Yang Yuan, Haonan Zhang, Xin Wang وآخرون · 2026

Logging code supports debugging, monitoring, and software maintenance, but excessive logging can add noise, impose runtime overhead, and obscure diagnostic information. While prior research has extensively studied logging code generation and modification, logging removal remains comparatively underexplored. In this pap …

نسخة أولية وصول مفتوح

ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing

Miaobo Hu, Shuhao Hu, Xiaobo Guo وآخرون · 2026

Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request supp …

نسخة أولية وصول مفتوح

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

Miaobo Hu, Shuhao Hu, Xiaobo Guo وآخرون · 2026

Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate t …

نسخة أولية وصول مفتوح

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

Eray Turkel, Mengsha Sun, Kartik Ayyar وآخرون · 2026

We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benc …

نسخة أولية وصول مفتوح

FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training

Miaobo Hu, Shuhao Hu, Xiaobo Guo وآخرون · 2026

Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its tra …

نسخة أولية وصول مفتوح

R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

Xin Wang, Zichuan Ying, Xinna Lin وآخرون · 2026

Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, e …

نسخة أولية وصول مفتوح

Multi-agent discussion gains less when dissent is withheld

Chand Sahil Mansuri, Xin Wang, Mengying Li وآخرون · 2026

Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it …

نسخة أولية وصول مفتوح

VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers

Miaobo Hu, Shuhao Hu, Xiaobo Guo وآخرون · 2026

Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertaint …

المؤلفون المشاركون