Login / Signup

Comparative Evaluation of LLMs in Clinical Oncology.

Nicholas R RydzewskiDeepak DinakaranShuang G ZhaoEytan RuppinIsmail Baris TurkbeyDeborah E CitrinKrishnan R Patel
Published in: NEJM AI (2024)
Of the models tested on a standardized set of oncology questions, GPT-4 was observed to have the highest performance. Although this performance is impressive, all LLMs continue to have clinically significant error rates, including examples of overconfidence and consistent inaccuracies. Given the enthusiasm to integrate these new implementations of AI into clinical practice, continued standardized evaluations of the strengths and limitations of these products will be critical to guide both patients and medical professionals. (Funded by the National Institutes of Health Clinical Center for Research and the Intramural Research Program of the National Institutes of Health; Z99 CA999999.).
Keyphrases