E-ISSN: 2587-0351 | ISSN: 1300-2694
Diagnostic Performance of Large Language Models Using FDG PET-Derived Z-Score Profiles and Structured Clinical Information in Neurodegenerative Diseases [Van Med J]
Van Med J. 2026; 33(3): 277-283 | DOI: 10.5505/vmj.2026.48303

Diagnostic Performance of Large Language Models Using FDG PET-Derived Z-Score Profiles and Structured Clinical Information in Neurodegenerative Diseases

Hasan Onner1, Lutfu Perktas1, Kevser Oksuzoglu1, Farise Yilmaz1, Fettah Eren2, Serefnur Ozturk2
1Selçuk University, Faculty Of Medicine, Department of Nuclear Medicine
2Selçuk University, Faculty of Medicine, Department of Neurology

INTRODUCTION: To evaluate the diagnostic performance of large language models (LLMs) with FDG PET-derived Z-score profiles and structured clinical information in patients with suspected neurodegenerative diseases (NDs).
METHODS: This retrospective study included patients who underwent FDG-PET imaging of the brain for suspected NDs. FDG PET brain imaging Z-score derived from a database and anonymized structured clinical information were provided to four LLMs (ChatGPT, Grok, Gemini, and DeepSeek). Each model generated a single diagnosis among Alzheimer’s disease, frontotemporal dementia, dementia with Lewy bodies, vascular dementia, primary progressive aphasia, or normal/nonspecific. A multidisciplinary consensus diagnosis served as the reference standard. LLM outputs were stratified into subgroups: overall, high diagnostic confidence (HC; ≥85%), epicenter concordance (EC), and combined (HC + EC). Diagnostic agreement was assessed using Cohen’s kappa (κ).
RESULTS: A total of 80 patients (42 females, 38 males) were included. All LLMs showed significant agreement with the diagnosis. ChatGPT had the highest agreement (κ=0.760), followed by Grok (κ=0.648) and DeepSeek (κ=0.639). In the HC subgroup, agreement improved across all models, with ChatGPT reaching κ=0.872, followed by DeepSeek (κ=0.789), Grok (κ=0.711), and Gemini (κ=0.676). In the EC subgroup, ChatGPT (κ=0.828) and Grok (κ=0.799) show substantial concordance, and DeepSeek (κ=0.682) and Gemini (κ=0.542) demonstrate moderate agreement. In the combined HC + EC subgroup, ChatGPT achieved the strongest performance (κ=0.860), followed by Grok (κ=0.824) and DeepSeek (κ=0.786), while Gemini reached moderate agreement (κ=0.652).
DISCUSSION AND CONCLUSION: LLMs showed moderate-to-high agreement with multidisciplinary consensus diagnoses in patients with suspected NDs. Agreement was higher in cases with high diagnostic confidence and epicenter concordance.

Keywords: fluorodeoxyglucose positron emission tomography, brain imaging, neurodegenerative diseases, large language models, artificial intelligence, decision support systems.


Corresponding Author: Hasan Onner, Türkiye
Manuscript Language: English
×
APA
NLM
AMA
MLA
Chicago
Copied!
CITE
LookUs & Online Makale