الباحثون

Xiang Xu

المنشورات 5

نسخة أولية وصول مفتوح

ALoDLM: Adaptively Looped Diffusion Language Models

Liancheng Fang, Zhuowei Li, Youngeun Kim وآخرون · 2026

Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed seq …

نسخة أولية وصول مفتوح

STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking

Wan Tian, Zhongyi Li, Xiang Xu وآخرون · 2026

Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseli …

نسخة أولية وصول مفتوح

Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

Zhongyi Li, Wan Tian, Xiang Xu وآخرون · 2026

Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weig …

نسخة أولية وصول مفتوح

FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

Xi Xiao, Yunbei Zhang, Chen Liu وآخرون · 2026

In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token …

نسخة أولية وصول مفتوح

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Xi Xiao, Tianchen Zhao, Youngeun Kim وآخرون · 2026

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent toke …

المؤلفون المشاركون