Abstract
Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we propose AHMAD, a simple yet effective framework for generalist multitask learning that integrates different key vision tasks: semantic segmentation, instance segmentation, depth estimation, keypoint detection, and object detection. Our approach incorporates these five tasks into a unified structure: a shared encoder-decoder with several lightweight task-specific projectors. Under the multitask learning paradigm, we observed a complementary performance gain, achieving a state-of-the-art PQ of 53.1 and an mIoU of 66.5 for COCO-val panoptic and semantic segmentation, respectively. Additionally, for top-down keypoint detection, which typically incurs high computational overhead due to multiple forward passes, we introduce a knowledge distillation-based method that enables a single forward pass over the entire image, greatly improving efficiency. Ultimately, our model delivers a lightweight yet effective generalist multitask learning framework, demonstrating strong performance across five vision tasks.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Mahdi, M., Prisadnikov, N., Fu, Y., Scribano, C., Paudel, D. P., & Van Gool, L. (2026). AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection. https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection
MLA 9
Mahdi, Mohammad, et al. "AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection." https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection.
Chicago (author–date)
Mahdi, Mohammad, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, and Luc Van Gool. 2026. "AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection." https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection.
Harvard
Mahdi, M., Prisadnikov, N., Fu, Y., Scribano, C., Paudel, D. P. and Van Gool, L. (2026) 'AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection', Available at: https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection.
Vancouver
Mahdi M, Prisadnikov N, Fu Y, Scribano C, Paudel DP, Van Gool L. AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection. https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection
IEEE
M. Mahdi, N. Prisadnikov, Y. Fu, C. Scribano, D. P. Paudel, and L. Van Gool, "AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection," https://omanscience.com/en/articles/ahmad-adaptive-hybrid-multi-task-vision-learning-with-assisted-distillation-for-keypoint-detection.