You Are Your Best Teacher: Semi-Supervised Surgical Point Tracking with Cycle-Consistent Self-Distillation
By: Valay Bundele, Mehran Hosseinzadeh, Hendrik Lensch
Potential Business Impact:
Helps robots track tiny things in messy videos.
Synthetic datasets have enabled significant progress in point tracking by providing large-scale, densely annotated supervision. However, deploying these models in real-world domains remains challenging due to domain shift and lack of labeled data-issues that are especially severe in surgical videos, where scenes exhibit complex tissue deformation, occlusion, and lighting variation. While recent approaches adapt synthetic-trained trackers to natural videos using teacher ensembles or augmentation-heavy pseudo-labeling pipelines, their effectiveness in high-shift domains like surgery remains unexplored. This work presents SurgTracker, a semi-supervised framework for adapting synthetic-trained point trackers to surgical video using filtered self-distillation. Pseudo-labels are generated online by a fixed teacher-identical in architecture and initialization to the student-and are filtered using a cycle consistency constraint to discard temporally inconsistent trajectories. This simple yet effective design enforces geometric consistency and provides stable supervision throughout training, without the computational overhead of maintaining multiple teachers. Experiments on the STIR benchmark show that SurgTracker improves tracking performance using only 80 unlabeled videos, demonstrating its potential for robust adaptation in high-shift, data-scarce domains.
Similar Papers
SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
CV and Pattern Recognition
Helps surgeons by automatically tracking surgery steps.
Dual Invariance Self-training for Reliable Semi-supervised Surgical Phase Recognition
Image and Video Processing
Helps surgeons by showing them what step they're on.
TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student Learning
CV and Pattern Recognition
Helps surgery robots see depth in moving video.