Abstract
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34\,cm and increases temporal mean average precision for motion segmentation from 44.54\% to 54.44\%.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Turkcan, M. K., Samal, S., & Kostic, Z. (2026). Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints. https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints
MLA 9
Turkcan, Mehmet Kerem, et al. "Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints." https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints.
Chicago (author–date)
Turkcan, Mehmet Kerem, Soham Samal, and Zoran Kostic. 2026. "Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints." https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints.
Harvard
Turkcan, M. K., Samal, S. and Kostic, Z. (2026) 'Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints', Available at: https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints.
Vancouver
Turkcan MK, Samal S, Kostic Z. Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints. https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints
IEEE
M. K. Turkcan, S. Samal, and Z. Kostic, "Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints," https://omanscience.com/en/articles/surgical-kinematics-from-monocular-video-with-learned-articulated-motion-constraints.