Surgical AI is trained under a permanent shortage of data. Privacy law prevents collecting it in bulk, and to combat it SOTA papers propose three solutions: pretrain foundation models on whatever can be pooled, aggregate across institutions, and lately generating the missing data with world models. These results were achieved by both pixel-space VLM models as well as V-JEPA based approaches. While those methods achieve ever increasing scores in detection, phase recognition and prediction, there is little research on the robustness and behavior of those models under unexpected conditions. What happens if the procedure changes mid-surgery due to an unexpected complication? Would it have been better to use procedure B instead of procedure A?
We want to analyze and compare how those models are able to hold up under those assumptions and define a metric that can be used for future projects. We hope that by having a clear baseline and the insights gained by this investigation, we'll be able to find new ways to guide current and future methods into having an enhanced understanding of its surroundings.
Eventually, this will hopefully lead to models that are able to accurately predict what would happen if a surgeon continued with their current action, propose different, safer approaches during a surgery, or even be used as a fully virtual training environment.