Comparison of the problem-solving performance of ChatGPT-3.5, ChatGPT-4, Bing Chat, and Bard for the Korean emergency medicine board examination question bank

Copyright © 2024 the Author(s). Published by Wolters Kluwer Health, Inc..

Large language models (LLMs) have been deployed in diverse fields, and the potential for their application in medicine has been explored through numerous studies. This study aimed to evaluate and compare the performance of ChatGPT-3.5, ChatGPT-4, Bing Chat, and Bard for the Emergency Medicine Board Examination question bank in the Korean language. Of the 2353 questions in the question bank, 150 questions were randomly selected, and 27 containing figures were excluded. Questions that required abilities such as analysis, creative thinking, evaluation, and synthesis were classified as higher-order questions, and those that required only recall, memory, and factual information in response were classified as lower-order questions. The answers and explanations obtained by inputting the 123 questions into the LLMs were analyzed and compared. ChatGPT-4 (75.6%) and Bing Chat (70.7%) showed higher correct response rates than ChatGPT-3.5 (56.9%) and Bard (51.2%). ChatGPT-4 showed the highest correct response rate for the higher-order questions at 76.5%, and Bard and Bing Chat showed the highest rate for the lower-order questions at 71.4%. The appropriateness of the explanation for the answer was significantly higher for ChatGPT-4 and Bing Chat than for ChatGPT-3.5 and Bard (75.6%, 68.3%, 52.8%, and 50.4%, respectively). ChatGPT-4 and Bing Chat outperformed ChatGPT-3.5 and Bard in answering a random selection of Emergency Medicine Board Examination questions in the Korean language.

Medienart:

E-Artikel

Erscheinungsjahr:

2024

Erschienen:

2024

Enthalten in:

Zur Gesamtaufnahme - volume:103

Enthalten in:

Medicine - 103(2024), 9 vom: 01. März, Seite e37325

Sprache:

Englisch

Beteiligte Personen:

Lee, Go Un [VerfasserIn]
Hong, Dae Young [VerfasserIn]
Kim, Sin Young [VerfasserIn]
Kim, Jong Won [VerfasserIn]
Lee, Young Hwan [VerfasserIn]
Park, Sang O [VerfasserIn]
Lee, Kyeong Ryong [VerfasserIn]

Links:

Volltext

Themen:

Comparative Study
Journal Article

Anmerkungen:

Date Completed 04.03.2024

Date Revised 04.03.2024

published: Print

Citation Status MEDLINE

doi:

10.1097/MD.0000000000037325

funding:

Förderinstitution / Projekttitel:

PPN (Katalog-ID):

NLM369187709