الباحثون

Min Zhang

المنشورات 10

نسخة أولية وصول مفتوح

Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes

Hao Li, Jinye Zhang, Bobo Li وآخرون · 2026

Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE …

نسخة أولية وصول مفتوح

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

Hexuan Deng, Zihao Yan, Xuebo Liu وآخرون · 2026

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a …

نسخة أولية وصول مفتوح

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

Zipeng Wang, Xinpeng Dong, Yuefan Wang وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not lear …

نسخة أولية وصول مفتوح

ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning

Bingchen Pei, Lichong Chen, Bingxi Zhao وآخرون · 2026

Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

نسخة أولية وصول مفتوح

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

Yilun Liu, Shimin Tao, Minggui He وآخرون · 2026

Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However, this ecosystem remains deeply English-centric: ou …

نسخة أولية وصول مفتوح

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

Kunwei Wu, Xiang Liu, Guocai Yao وآخرون · 2026

Counterfactual world models (CWM) extract motion from pretrained video predictors by comparing factual and intervened predictions, but uniform aggregation weights responses equally without explicitly incorporating physical priors. Our key insight is to incorporate physical priors into candidate reliability learning, mo …

المؤلفون المشاركون