Volver a resultados
Ficha bibliográfica · Consulta y acceso
Artículo

Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine

Zehua Jiang et al · Elsevier · 2026

Acceso abierto disponible
Lectura rápida. Revisá los datos básicos del recurso y luego accedé al contenido desde el botón principal. En esta ficha solo se muestra la información necesaria para identificar la obra, citarla y abrirla.

Acceso al recurso

Entrá al contenido desde la opción principal o elegí otra fuente disponible.

DOAJ DOAJ Articles
Entrar por DOAJ
Acceso principal

Acceso abierto disponible

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Abrir recurso

Resumen

Descripción general del contenido del recurso.

Large language models (LLMs) have demonstrated encouraging performance for medical natural language processing (NLP) tasks, approaching human-equivalent performance in some of the standard benchmarks, positioning them as game-changers in healthcare. However, there remains a persistent gap between high benchmark performance and clinical utility of NLP algorithms due to limitations of existing evaluation paradigms. Existing benchmarks tend to use static, task-specific benchmarks, and as a result, they do not capture the full dimension of complexity, safety, interpretability, and integration in workflow required for the safe deployment in the clinic. The editorial advocates for a shift in evaluation paradigms from narrow score-based metrics to a dynamic, multi-dimensional, and patient-centered system of clinical gatekeeping. The proposed framework integrates a four-phase process, including retrospective benchmarking, pilot testing, multi-center validation, and real-world monitoring, alongside a capability-task-behavior-value progression and a continuous human-in-the-loop feedback mechanism. This comprehensive strategy ensures not only technical robustness but also clinical relevance, ethical accountability, and adaptive improvement, transforming LLMs from experimental tools into reliable clinical partners for safer and more patient-centric healthcare delivery.

Cómo citar

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Z. J. E. (2026). Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine. https://doi.org/10.1016/j.imed.2026.01.001

MLA

al, Zehua Jiang et. "Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine." 2026. https://doi.org/10.1016/j.imed.2026.01.001.

Chicago

al, Zehua Jiang et. 2026. "Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine.". https://doi.org/10.1016/j.imed.2026.01.001.

Harvard

al, Z. J. E. 2026, Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine, Elsevier, available at: https://doi.org/10.1016/j.imed.2026.01.001 [Accessed 8 Aug. 2026].

Compartir e imprimir

Guardá la ficha, copiá su enlace permanente o imprimila como PDF.

Exportar referencia

Si usás un gestor bibliográfico, podés exportar el registro en los formatos más comunes.

Detalles del recurso

Información bibliográfica útil para confirmar que se trata del material correcto.

Título
Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine
Autor / colaboradores
Zehua Jiang et al
Editorial
Elsevier
Año de publicación
2026
ISSN
2667-1026
ISSN
2667-1026
Idioma
Inglés

Materias

Explorá otros recursos relacionados a partir de estas materias.

Copiado