Details der Publikation - Reliability of large language models in managing odontogenic sinusitis clinical scenarios

Reliability of large language models in managing odontogenic sinusitis clinical scenarios : a preliminary multidisciplinary evaluation

© 2024. The Author(s)..

PURPOSE: This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifically for the multidisciplinary management of odontogenic sinusitis (ODS).

METHODS: A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS-related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators.

RESULTS: While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs' responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators' specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs.

CONCLUSIONS: LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs' performance as they evolve over time.

Medienart:	E-Artikel

Erscheinungsjahr:	2024
Erschienen:	2024

Enthalten in:	Zur Gesamtaufnahme - volume:281
Enthalten in:	European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery - 281(2024), 4 vom: 18. Apr., Seite 1835-1841

Sprache:	Englisch

Beteiligte Personen:	Saibene, Alberto Maria [VerfasserIn] Allevi, Fabiana [VerfasserIn] Calvo-Henriquez, Christian [VerfasserIn] Maniaci, Antonino [VerfasserIn] Mayo-Yáñez, Miguel [VerfasserIn] Paderno, Alberto [VerfasserIn] Vaira, Luigi Angelo [VerfasserIn] Felisati, Giovanni [VerfasserIn] Craig, John R [VerfasserIn]

Links:	Volltext

Themen:	Artificial intelligence Chronic rhinosinusitis Computer-assisted diagnosis Dental implant Journal Article Maxillary sinusitis Oroantral fistula

Anmerkungen:	Date Completed 18.03.2024 Date Revised 02.04.2024 published: Print-Electronic Citation Status MEDLINE

doi:	10.1007/s00405-023-08372-4

funding:
Förderinstitution / Projekttitel:

PPN (Katalog-ID):	NLM366805959

Internformat


LEADER	01000caa a22002652 4500
001	NLM366805959
003	DE-627
005	20240403000036.0
007	cr uuu---uuuuu
008	240114s2024 xx \|\|\|\|\|o 00\| \|\|eng c
024	7		\|a 10.1007/s00405-023-08372-4 \|2 doi
028	5	2	\|a pubmed24n1361.xml
035			\|a (DE-627)NLM366805959
035			\|a (NLM)38189967
040			\|a DE-627 \|b ger \|c DE-627 \|e rakwb
041			\|a eng
100	1		\|a Saibene, Alberto Maria \|e verfasserin \|4 aut
245	1	0	\|a Reliability of large language models in managing odontogenic sinusitis clinical scenarios \|b a preliminary multidisciplinary evaluation
264		1	\|c 2024
336			\|a Text \|b txt \|2 rdacontent
337			\|a ƒaComputermedien \|b c \|2 rdamedia
338			\|a ƒa Online-Ressource \|b cr \|2 rdacarrier
500			\|a Date Completed 18.03.2024
500			\|a Date Revised 02.04.2024
500			\|a published: Print-Electronic
500			\|a Citation Status MEDLINE
520			\|a © 2024. The Author(s).
520			\|a PURPOSE: This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifically for the multidisciplinary management of odontogenic sinusitis (ODS)
520			\|a METHODS: A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS-related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators
520			\|a RESULTS: While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs' responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators' specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs
520			\|a CONCLUSIONS: LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs' performance as they evolve over time
650		4	\|a Journal Article
650		4	\|a Artificial intelligence
650		4	\|a Chronic rhinosinusitis
650		4	\|a Computer-assisted diagnosis
650		4	\|a Dental implant
650		4	\|a Maxillary sinusitis
650		4	\|a Oroantral fistula
700	1		\|a Allevi, Fabiana \|e verfasserin \|4 aut
700	1		\|a Calvo-Henriquez, Christian \|e verfasserin \|4 aut
700	1		\|a Maniaci, Antonino \|e verfasserin \|4 aut
700	1		\|a Mayo-Yáñez, Miguel \|e verfasserin \|4 aut
700	1		\|a Paderno, Alberto \|e verfasserin \|4 aut
700	1		\|a Vaira, Luigi Angelo \|e verfasserin \|4 aut
700	1		\|a Felisati, Giovanni \|e verfasserin \|4 aut
700	1		\|a Craig, John R \|e verfasserin \|4 aut
773	0	8	\|i Enthalten in \|t European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery \|d 1994 \|g 281(2024), 4 vom: 18. Apr., Seite 1835-1841 \|w (DE-627)NLM012637548 \|x 1434-4726 \|7 nnns
773	1	8	\|g volume:281 \|g year:2024 \|g number:4 \|g day:18 \|g month:04 \|g pages:1835-1841
856	4	0	\|u http://dx.doi.org/10.1007/s00405-023-08372-4 \|3 Volltext
912			\|a GBV_USEFLAG_A
912			\|a GBV_NLM
951			\|a AR
952			\|d 281 \|j 2024 \|e 4 \|b 18 \|c 04 \|h 1835-1841

Reliability of large language models in managing odontogenic sinusitis clinical scenarios : a preliminary multidisciplinary evaluation

Zugang & Verfügbarkeit

Zugehörige Publikationen/Bände