الباحثون

Minghua He

المنشورات 3

نسخة أولية وصول مفتوح

Playing social deduction games with reinforcement fine-tuned large language models

Lingzhe Zhang, Yunpeng Zhai, Tong Jia وآخرون · 2026

Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state infere …

نسخة أولية وصول مفتوح

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping وآخرون · 2026

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply …

نسخة أولية وصول مفتوح

GrowMTP: Can RL Grow Its Own Draft Head?

Minghua He, Lingzhe Zhang, Yuan Liu وآخرون · 2026

Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing subs …

المؤلفون المشاركون