الباحثون

Rui Qian

المنشورات 9

نسخة أولية وصول مفتوح

Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

Yi Wang, Rui Qian, Yu Li وآخرون · 2026

Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly …

نسخة أولية وصول مفتوح

MemoWM: How World Models Change What Agents Need to Remember

Bingfan Zeng, Zhisheng Chen, Chenbo Sang وآخرون · 2026

Long-term agents face growing storage demands as they accumulate experience. World models capture reusable regularities that can reduce the information stored for each experience. We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to …

نسخة أولية وصول مفتوح

UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting

Jinshi Liu, Pan Liu, Lei He وآخرون · 2026

Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given …

نسخة أولية وصول مفتوح

Reliability-Aware Checkpoint Selection for Domain Generalization

Jinshi Liu, Jiahao Li, Pan Liu وآخرون · 2026

Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We …

نسخة أولية وصول مفتوح

Learning to Reason with Persistent Object States for Video Instance Segmentation

Yongxue Xu, Boxue Yang, Ziqian Liu وآخرون · 2026

Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, …

نسخة أولية وصول مفتوح

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

Xutao Mao, Rui Qian, Longxiang Wang وآخرون · 2026

LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, and destructive over-c …

نسخة أولية وصول مفتوح

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

Yuhe Wu, Rui Qian, Guangyu Wang وآخرون · 2026

Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be o …

المؤلفون المشاركون