Abstract
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Shen, L., Tang, W., Karray, F., & Zhang, M. L. (2026). Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning. https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning
MLA 9
Shen, Lingyu, et al. "Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning." https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning.
Chicago (author–date)
Shen, Lingyu, Wei Tang, Fakhri Karray, and Min-Ling Zhang. 2026. "Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning." https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning.
Harvard
Shen, L., Tang, W., Karray, F. and Zhang, M. L. (2026) 'Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning', Available at: https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning.
Vancouver
Shen L, Tang W, Karray F, Zhang ML. Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning. https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning
IEEE
L. Shen, W. Tang, F. Karray, and M. L. Zhang, "Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning," https://omanscience.com/en/articles/joint-class-time-learning-for-video-classification-with-multi-instance-partial-label-learning.