E-ISSN: 2587-0351 | ISSN: 1300-2694
Can Artificial Intelligence Calculate Disability Rates? A Comparative Study of Three Large Language Models Against Expert Assessment [Van Med J]
Van Med J. 2026; 33(3): 228-237 | DOI: 10.5505/vmj.2026.93899

Can Artificial Intelligence Calculate Disability Rates? A Comparative Study of Three Large Language Models Against Expert Assessment

Yasin Etli1, Erhan Kartal1, Mahmut Asirdizer2
1Van Yüzüncü Yıl University, Faculty of Medicine, Department of Forensic Medicine
2Bahçeşehir University, Faculty of Medicine, Department of Forensic Medicine

INTRODUCTION: Disability impairment rating is a critical medicolegal function requiring precise integration of clinical findings with regulatory tables. Large language models (LLMs) have shown promising capabilities in medicine; however, their ability to perform quantitative disability assessments has not been evaluated. This study aimed to assess the agreement between three commercial LLMs and expert-determined disability ratings.
METHODS: One hundred standardized fictional clinical cases were constructed representing the spectrum of medicolegal disability assessments. Each case was evaluated by a forensic medicine specialist using the Turkish Regulation on Disability Assessment for Adults (2019) as the reference standard. Three LLMs—GPT-5, Gemini 3 Pro, and Claude Opus 4.6—were tested using a zero-shot prompt. A three-model ensemble was also examined. Agreement was assessed using intraclass correlation coefficients (ICC), Bland-Altman analysis, and mean absolute error (MAE).
RESULTS: All three LLMs demonstrated good agreement with the reference standard, with ICC values of 0.875 (Gemini 3 Pro), 0.848 (Claude Opus 4.6), and 0.814 (GPT-5). MAE ranged from 5.1 to 6.0 percentage points. No significant difference was found among models (Friedman p=0.790). The three-model ensemble achieved the highest ICC (0.890) and lowest MAE (4.7 points). Accuracy declined substantially for rates ≥21% (MAE 9.1–11.5). Agreement for temporary work disability duration was moderate (ICC 0.701–0.750), while caregiver need showed poor-to-moderate agreement (ICC 0.497–0.585).
DISCUSSION AND CONCLUSION: Commercial LLMs demonstrate good overall agreement with expert disability ratings but show declining accuracy for complex multi-system cases. A multi-model ensemble approach improves reliability. LLMs may serve as supportive tools in medicolegal disability assessment, although expert oversight remains essential.

Keywords: Artificial intelligence, Disability evaluation, Forensic medicine, Large language models, Reproducibility of results, Medicolegal assessment


Corresponding Author: Yasin Etli, Türkiye
Manuscript Language: English
×
APA
NLM
AMA
MLA
Chicago
Copied!
CITE
LookUs & Online Makale