نسخة أولية وصول مفتوح
RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, …