Can AI Think Like a Plastic Surgeon? Evaluating GPT-4's Clinical Judgment in Reconstructive Procedures of the Upper Extremity

被引：7

作者：

Leypold, Tim ^{[1
]}

Schaefer, Benedikt ^{[1
]}

Boos, Anja ^{[1
]}

Beier, Justus P. ^{[1
]}

机构：

[1] Univ Hosp RWTH Aachen, Hand Surg Burn Ctr, Dept Plast Surg, Aachen, Germany

来源：

PLASTIC AND RECONSTRUCTIVE SURGERY-GLOBAL OPEN | 2023年 / 11卷 / 12期

关键词：

D O I：

10.1097/GOX.0000000000005471

中图分类号：

R61 [外科手术学];

学科分类号：

摘要：

This study delves into the potential application of OpenAI's Generative Pretrained Transformer 4 (GPT-4) in plastic surgery, with a particular focus on procedures involving the hand and arm. GPT-4, a cutting-edge artificial intelligence (AI) model known for its advanced chat interface, was tested on nine surgical scenarios of varying complexity. To optimize the performance of GPT-4, prompt engineering techniques were used to guide the model's responses and improve the relevance and accuracy of its output. A panel of expert plastic surgeons evaluated the responses using a Likert scale to assess the model's performance, based on five distinct criteria. Each criterion was scored on a scale of 1 to 5, with 5 representing the highest possible score. GPT-4 demonstrated a high level of performance, achieving an average score of 4.34 across all cases, consistent across different complexities. The study highlights the ability of GPT-4 to understand and respond to complicated surgical scenarios. However, the study also identifies potential areas for improvement. These include refining the prompts used to elicit responses from the model and providing targeted training with specialized, up-to-date sources. This study demonstrates a new approach to exploring large language models and highlights potential future applications of AI. These could improve patient care, refine surgical outcomes, and even change the way we approach complex clinical scenarios in plastic surgery. However, the intrinsic limitations of AI in its current state, together with the potential ethical considerations and the inherent uncertainty of unanticipated issues, serve to reiterate the indispensable role and unparalleled value of human plastic surgeons.

引用

页数：3

共 2 条

[1] Evaluating prompt engineering on GPT-3.5's performance in USMLE-style medical calculations and clinical scenarios generated by GPT-4
Patel, Dhavalkumar
Raut, Ganesh
Zimlichman, Eyal
Cheetirala, Satya Narayan
Nadkarni, Girish N.
Glicksberg, Benjamin S.
Apakama, Donald U.
Bell, Elijah J.
Freeman, Robert
Timsina, Prem
Klang, Eyal
[J]. SCIENTIFIC REPORTS, 2024, 14 (01):
[2] Can large language models replace humans in systematic reviews? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
Khraisha, Qusai
Put, Sophie
Kappenberg, Johanna
Warraitch, Azza
Hadfield, Kristin
[J]. RESEARCH SYNTHESIS METHODS, 2024, 15 (04) : 616 - 626

← 1 →