الباحثون

Min-Ling Zhang

المنشورات 4

نسخة أولية وصول مفتوح

Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

Lingyu Shen, Wei Tang, Fakhri Karray وآخرون · 2026

Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples labe …

نسخة أولية وصول مفتوح

Outcome-Guided On-Policy Self-Distillation

ZheXu Wang, Mao-Lin Luo, Yankun Hong وآخرون · 2026

On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-toke …

نسخة أولية وصول مفتوح

Learning Process Rewards via Reasoning State Propagation

Kai Gan, Zi-Hao Zhou, Bo Ye وآخرون · 2026

Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complemen …

المؤلفون المشاركون