Abstract
Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S$^3$KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhou, K., Chen, C., Zeng, R., Dai, M., Yang, Y., Hu, J., Li, D., Tan, M., & Liu, F. (2026). Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study. https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study
MLA 9
Zhou, Kai, et al. "Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study." https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study.
Chicago (author–date)
Zhou, Kai, Chuanshen Chen, Runhao Zeng, Meng Dai, Yifan Yang, Jinwu Hu, Daiyuan Li, Mingkui Tan, and Fei Liu. 2026. "Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study." https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study.
Harvard
Zhou, K., Chen, C., Zeng, R., Dai, M., Yang, Y., Hu, J., Li, D., Tan, M. and Liu, F. (2026) 'Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study', Available at: https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study.
Vancouver
Zhou K, Chen C, Zeng R, Dai M, Yang Y, Hu J, et al. Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study. https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study
IEEE
K. Zhou, C. Chen, R. Zeng, M. Dai, Y. Yang, J. Hu, D. Li, M. Tan, and F. Liu, "Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study," https://omanscience.com/en/articles/structure-aware-keypoint-localization-for-videofluoroscopic-swallowing-study.