SurgGoal: Rethinking Surgical Planning Evaluation via Goal-Satisfiability
By: Ruochen Li , Kun Yuan , Yufei Xia and more
Surgical planning integrates visual perception, long-horizon reasoning, and procedural knowledge, yet it remains unclear whether current evaluation protocols reliably assess vision-language models (VLMs) in safety-critical settings. Motivated by a goal-oriented view of surgical planning, we define planning correctness via phase-goal satisfiability, where plan validity is determined by expert-defined surgical rules. Based on this definition, we introduce a multicentric meta-evaluation benchmark with valid procedural variations and invalid plans containing order and content errors. Using this benchmark, we show that sequence similarity metrics systematically misjudge planning quality, penalizing valid plans while failing to identify invalid ones. We therefore adopt a rule-based goal-satisfiability metric as a high-precision meta-evaluation reference to assess Video-LLMs under progressively constrained settings, revealing failures due to perception errors and under-constrained reasoning. Structural knowledge consistently improves performance, whereas semantic guidance alone is unreliable and benefits larger models only when combined with structural constraints.
Similar Papers
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
CV and Pattern Recognition
Helps surgeons by understanding surgery videos.
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
CV and Pattern Recognition
Helps robots learn to fix their own mistakes.
SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
CV and Pattern Recognition
Helps robot surgeons see and understand actions.