Does liver-surgery AI work beyond the hospital where it was trained?
- Publication state
- PublishedPublished after author review.
- Evidence basis
- 2Connected public source records listed below.
The study tested this question on eight videos from two hospitals outside the development center. That is an initial test of transferability, not proof that the system will work reliably everywhere.
The model in this 2025 study, coauthored by Namkee Oh, labeled vascular structures and the avascular plane during right liver mobilization in laparoscopic donor surgery. Forty videos from Samsung Medical Center supported patient-level five-fold cross-validation. Eight videos from Myongji Hospital and Yeungnam University Medical Center were kept out of model development and used for external validation.
Why separate the hospitals? Testing on new patients at the development center asks whether the model can handle additional cases in that setting. Testing on data from other institutions asks a different question: does its performance carry beyond the setting in which it was developed?
The reported average Dice scores—a measure of overlap between predicted regions and human annotations—were numerically similar in internal and external evaluation. But the study did not calculate confidence intervals or establish statistical equivalence between those groups. A small external test cannot show how the model will behave across the full variety of operating rooms.
The practical lesson is to look beyond the phrase “multicenter study” and ask which hospitals supplied training data, which supplied test data, and what remained untested. Here, the evidence concerns retrospective donor-surgery videos. Whether using the system reduces bleeding or vascular injury still requires prospective clinical evaluation.

간 수술 AI는 학습한 병원 밖에서도 작동할까요?
이 연구는 모델 개발에 사용하지 않은 다른 두 병원의 영상 8건으로 이 질문을 시험했습니다. 다른 환경으로 성능이 이어지는지 살펴본 초기 검증이지, 어디서나 안정적으로 작동한다는 증명은 아닙니다.
오남기가 공동저자로 참여한 이 2025년 연구의 모델은 복강경 간 공여자 수술의 우간 가동화 과정에서 혈관 구조와 무혈관 박리면을 표시했습니다. 삼성서울병원 영상 40건으로 환자 단위의 5겹 교차검증을 수행했습니다. 명지병원과 영남대학교병원의 영상 8건은 모델 개발에 사용하지 않고 외부 검증에 사용했습니다.
왜 병원을 구분할까요? 개발 병원의 새로운 환자 자료로 시험하는 것은 그 환경에서 다른 사례도 다룰 수 있는지 묻는 일입니다. 다른 기관의 자료로 시험하는 것은 별개의 질문입니다. 개발된 환경을 벗어나서도 성능이 이어질까요?
예측 영역과 사람이 표시한 영역이 얼마나 겹치는지를 나타내는 평균 Dice 점수는 내부·외부 평가에서 수치상 비슷했습니다. 그러나 이 연구는 신뢰구간을 계산하거나 두 집단에서 성능이 통계적으로 동등함을 입증하지는 않았습니다. 소규모 외부 검증만으로 다양한 수술실 환경 전반에서 모델이 어떻게 작동할지 알 수는 없습니다.
실용적인 교훈은 ‘다기관 연구’라는 표현에서 한 걸음 더 들어가, 어느 병원이 학습 자료를 제공했고 어느 병원이 검증 자료를 제공했으며 무엇이 아직 검증되지 않았는지 묻는 것입니다. 여기서 확보한 근거는 후향적 간 공여자 수술 영상에 관한 것입니다. 시스템 사용이 출혈이나 혈관 손상을 줄이는지는 앞으로 전향적인 임상 평가가 필요합니다.