Clinical Trial: Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes
Study Status: COMPLETED
Recruit Status: COMPLETED
Condition:
Emergency Medicine
Diagnostic Errors
Artificial Intelligence (AI) in Diagnosis
Study Type: OBSERVATIONAL
Official Title: Diagnostic Accuracy of Two Large Language Models Against a Blinded Specialist Consensus Standard in Turkish Emergency Department Notes: A Retrospective Study of 600 Cases
Brief Summary:
This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes.The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at…
Read more