Back to results
Bibliographic record · Consultation and access
Artículo

From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models

Nanziba Tasneem et al · Elsevier · 2026

Open access available
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

DOAJ DOAJ Articles
Entrar por DOAJ
Main access

Open access available

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Open resource

Summary

Descripción general del contenido del recurso.

Background: Large language models (LLMs), a revolutionary breakthrough in artificial intelligence, can be leveraged to automatically generate impressions for radiology reports, which usually require time, effort, and training. Our objective was to evaluate the performance of five recent LLMs (GPT–4, GPT–4o mini, Gemini 1.5–Pro, Gemini 1.5–Flash, and Llama 3.1) for impression generation. Methods: In this retrospective study, 100 radiology reports were sampled (20 from each of the report-groups 0–400, 400–800, 800–1,200, 1,200–2,000, and 2,000–8,000 based on character count of the Findings section) from the publicly available “BioNLP 2023 report summarization” dataset (collected between 2001–2016, training subset of size 59,320 considered for sampling), sourced from PhysioNet. Then, each of the five LLMs was zero-shot prompted to generate impressions using the findings from the sample. Generated impressions were evaluated: (a) subjectively for coherence, comprehensiveness, conciseness, and medical harmfulness by two radiology fellows and a large reasoning model (LRM) Gemini 2.5–Pro, and (b) objectively using a composite accuracy metric including recall-oriented understudy for gisting evaluation (ROUGE)-1, bilingual evaluation understudy (BLEU) and cosine similarity, against the original human expert-generated impressions. The LLMs were ranked according to the percentage agreement ranking of subjective and composite scores. Statistical tests ( Friedman and post-hoc Nemenyi tests) were used to assess inter-model differences. Results: The top-ranked models were Gemini 1.5–Pro, GPT–4, and Gemini 1.5–Flash. Performance varied across models for both human and LRM raters (Friedman test: Human P <1.82×10⁻⁶; LRM P <9.10×10⁻⁴⁰). Composite accuracy scores were significantly higher for the top three models (0.69, 0.68, and 0.68) versus others (0.65; Nemenyi P <1.11×10⁻¹⁶). The LRM aligned closely with human raters (2.15% complete disagreement) and identified all human-rated inaccurate impressions. Conclusion: Gemini 1.5–Pro outperformed GPT–4, in terms of coherence, comprehensiveness, and medical harmfulness, at a lower cost. Human and LRM evaluations were generally consistent, though the LRM was more conservative.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, N. T. E. (2026). From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models. https://doi.org/10.1016/j.imed.2025.11.003

MLA

al, Nanziba Tasneem et. "From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models." 2026. https://doi.org/10.1016/j.imed.2025.11.003.

Chicago

al, Nanziba Tasneem et. 2026. "From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models.". https://doi.org/10.1016/j.imed.2025.11.003.

Harvard

al, N. T. E. 2026, From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models, Elsevier, available at: https://doi.org/10.1016/j.imed.2025.11.003 [Accessed 9 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language models
Author / contributors
Nanziba Tasneem et al
Publisher
Elsevier
Publication year
2026
ISSN
2667-1026
ISSN
2667-1026
Language
English

Subjects

Explore related resources through these subjects.

Copied