Essay

When liver-surgery AI disagrees with its reference label, which is wrong?

Not every disagreement proves an AI error. Human annotations provide a comparison standard, but a mismatch also calls for review of the underlying surgical image.

Editorial cycle

From observation to a citable position.

Ideas move from operative observation to a reviewed and citable public position.

  1. 01ObserveClinical question
  2. 02ConnectPublished evidence
  3. 03WriteWorking thesis
  4. 04ReviewAuthor approval
  5. 05PublishStable record
Evidence boundaryPrimary-source records only
Publication indexGoogle Scholar Research identityORCID

When liver-surgery AI disagrees with its reference label, which is wrong?

Research insight · Right liver mobilization · 03 · Published

By Namkee Oh · 오남기

Publication state
PublishedPublished after author review.
Evidence basis
2Connected public source records listed below.

Not every disagreement proves an AI error. Human annotations provide a comparison standard, but a mismatch also calls for review of the underlying surgical image.

In image segmentation, a model outlines structures and its predictions are compared with human-drawn regions. This makes the reference annotation central to evaluation: a score describes agreement with that reference, not an independent verdict on anatomy.

The 2025 liver-mobilization study co-first-authored by Namkee Oh illustrates this distinction. Surgical staff labeled images from laparoscopic donor operations, and surgeons reviewed the annotations. Yet the authors presented examples in which model predictions included vessels absent from the reference masks. They interpreted these examples as evidence that annotation omissions could contribute to apparently incorrect predictions.

There is an important limit: the team did not relabel the full dataset and repeat the evaluation. These examples do not establish how often the model was right across all disagreements, or that it outperformed surgeons.

The practical implication, offered here as an editorial interpretation, is to make disagreements reviewable. Keep the original image, reference annotation, and prediction together. Examine the disputed region before revising either the label or the conclusion. A mismatch may reflect incomplete labeling, an incorrect prediction, or an uncertain boundary; the model alone cannot settle that question.

For liver-surgery research, scrutinizing the comparison standard belongs alongside scrutinizing the model. It does not replace clinical evaluation: this retrospective donor-video study did not test whether assistance reduced bleeding or vascular injury.

Three-panel conceptual diagram. Two panels show the same simplified scene with differing reference and predicted outlines around a small branching structure. A dashed ring marks the mismatch. The last panel asks reviewers to consider an incomplete label, prediction error, or uncertain boundary without declaring either outline correct.
Conceptual schematic of disagreement between a reference annotation and a model prediction in liver-surgery image analysis. The highlighted mismatch prompts review of the source image; it does not establish which output is correct. Original AI-generated conceptual illustration, not patient data, an actual model output, or a reproduction of the paper’s figures. Source: Oh et al., Scientific Reports (2025), DOI: 10.1038/s41598-025-11627-1.

한국어

간 수술 AI와 ‘정답 라벨’이 다르면, 어느 쪽이 틀린 걸까요?

불일치가 곧 AI의 오류를 뜻하는 것은 아닙니다. 사람이 만든 라벨은 비교 기준이지만, 예측이 라벨과 다를 때에는 원래 수술 영상도 함께 검토해야 합니다.

영상 분할에서는 모델이 구조물의 영역을 표시하고, 이를 사람이 그린 영역과 비교합니다. 따라서 기준 라벨은 평가의 핵심입니다. 점수는 그 기준과 얼마나 일치하는지를 나타내며, 해부학적 진실을 독립적으로 판정하는 것은 아닙니다.

오남기가 공동 제1저자로 참여한 2025년 우간 가동화 연구는 이 차이를 보여줍니다. 수술 인력이 복강경 공여자 수술 영상에 라벨을 만들고 외과의가 이를 검토했습니다. 그럼에도 연구진은 기준 라벨에 빠진 혈관을 모델이 표시한 사례들을 제시했습니다. 연구진은 이를 라벨의 누락이 겉보기에는 잘못된 예측으로 평가되는 데 영향을 줄 수 있다는 근거로 해석했습니다.

여기에는 중요한 한계가 있습니다. 전체 자료를 다시 라벨링한 뒤 평가를 반복하지는 않았습니다. 따라서 이 사례들만으로 모든 불일치 중 모델이 맞았던 비율을 알거나, 모델이 외과의보다 우수했다고 결론 내릴 수 없습니다.

여기서 제안하는 편집적 해석은 불일치를 검토할 수 있게 남기자는 것입니다. 원영상, 기준 라벨, 예측 결과를 함께 두고, 라벨이나 결론을 바꾸기 전에 서로 다른 부분을 살펴보는 것입니다. 불일치는 라벨의 누락, 잘못된 예측, 또는 모호한 경계에서 생길 수 있습니다. 모델만으로 어느 경우인지 결정할 수는 없습니다.

간 수술 AI를 연구할 때는 모델과 함께 비교 기준도 검토해야 합니다. 그렇다고 임상 평가를 대신할 수는 없습니다. 이 후향적 공여자 수술 영상 연구는 보조 시스템이 출혈이나 혈관 손상을 줄이는지 시험하지 않았습니다.

Connected public evidence