Objective This scoping review aimed to examine studies of nonagentic large language models (LLMs), LLM-based agents, and multiagent systems in emergency medicine and to identify current research trends and major gaps by analyzing their scope of clinical application, system structures, evaluation approaches, and input data characteristics.
Methods The Web of Science, Scopus, PubMed, and CINAHL were searched for literature published from March 8, 2021, to March 7, 2026. English-language full-text studies published during this period were included if they addressed the application, evaluation, or benchmarking of nonagentic LLMs, LLM-based agents, or multiagent systems in emergency medicine or the emergency department. Additional studies were identified through reference tracking and supplementary searching. In total, 35 studies were analyzed.
Results Of the 35 included studies, 26 were application studies, 6 were framework studies, and 3 were benchmark studies. Nonagentic LLMs were the most common system type (n=25), followed by LLM-based agents (n=7) and multiagent systems (n=3). Inputs were predominantly text-based, and evaluation mainly relied on expert comparison, retrospective record review, vignette-based comparison, and task-specific performance metrics. In contrast, workflow-level, prospective, and safety- and trustworthiness-oriented evaluations were limited.
Conclusion LLMs in emergency medicine have shown potential for task-level decision support and documentation. However, the current literature remains focused on nonagentic LLM-based task support, whereas studies that reflect the dynamic workflow of real emergency departments remain limited. Future research should expand toward workflow-aware design, operational evaluation, multimodal data integration, multiagent-based role coordination, and validation of safety and trustworthiness.
Citations
Citations to this article as recorded by
An Evidence-Based Framework for Patient-Facing Artificial Intelligence Integration in the Emergency Department Tehreem Rehman, Philip Jarrett, Joshua Lesko, James Augustine, Bradley D. Shy, Rohit B. Sangal, Nicholas Genes, Donald U. Apakama, Ethan E. Abbott, Abhi Mehrotra, Richard Andrew Taylor JACEP Open.2026; 7(5): 100466. CrossRef
Objective This study aimed to develop and validate MEDIVAL (Medical Documentation Validation), a progressive chain-of-thought (CoT) evaluation framework for automated assessment of large language model (LLM)-generated emergency department documentation, designed to align with expert clinical judgment in acute care settings. Methods We designed a three-tier evaluation framework incorporating persona-based, error-enhanced, and insight-integrated strategies. The framework was tested across four LLMs (GPT-4o, GPT-4.1, Claude-3.5, Claude-3.7) on 33 emergency department records reviewed by four expert emergency physicians. Each model applied the three CoT strategies across five criteria: appropriateness, accuracy, structure/format, conciseness, and clinical validity. Model outputs were compared with expert ratings using Spearman correlation coefficients. Differences were analyzed with the Friedman test and Wilcoxon signed rank test with Bonferroni correction. Reproducibility was assessed through intraclass correlation coefficient (ICC) analysis. Results All models demonstrated stronger alignment with expert ratings as CoT complexity increased, with Claude-3.7 (r=0.712, P<0.001) and GPT-4o (r=0.702, P<0.001) showing the highest correlations under the insight-integrated strategy. GPT-4.1 showed the greatest relative improvement (43.3% increase, r=0.457 to r=0.655, P<0.001). Significant overall differences were observed across strategies (χ2 (2)=48.39, P<0.001), though the error-enhanced and insight-integrated approaches differed only modestly yet significantly (P=0.002). High reproducibility was confirmed (ICC >0.919), with Claude-3.5 achieving the most consistent results (ICC, 0.997–0.998). Conclusion MEDIVAL demonstrates that progressive CoT strategies systematically improve automated evaluation of emergency department documentation while maintaining excellent reproducibility. This framework offers a viable prescreening tool to reduce expert workload and support reliable artificial intelligence integration into emergency medicine workflows.