Volver a resultados
Ficha bibliográfica · Consulta y acceso
Artículo

VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction

Yancheng Ling et al · Tsinghua University Press · 2026

Acceso abierto disponible
Lectura rápida. Revisá los datos básicos del recurso y luego accedé al contenido desde el botón principal. En esta ficha solo se muestra la información necesaria para identificar la obra, citarla y abrirla.

Acceso al recurso

Entrá al contenido desde la opción principal o elegí otra fuente disponible.

DOAJ DOAJ Articles
Entrar por DOAJ
Acceso principal

Acceso abierto disponible

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Abrir recurso

Resumen

Descripción general del contenido del recurso.

Pedestrian crossing intention prediction is crucial for autonomous driving. While existing models have achieved high accuracy, their generalization and robustness remain limited, hindering their performance in real-world scenarios. To overcome these limitations, we introduce the LVLMPed-CoT, a large vision language model (LVLM) that incorporates a chain-of-thought (CoT) mechanism to enhance pedestrian crossing intention prediction. It takes multimodal data as input and employs data distillation along with a two stage fine-tuning strategy to elicit the implicit CoT capability of a lightweight vision-language model for enhanced perception, reasoning, and prediction. The unified LVLMPed-CoT is trained on a joint open-source dataset (JAAD and PIE) and achieves superior or comparable performance to state-of-the-art models on both large-scale public datasets. The ablation study validates the contribution of the CoT prompt design and the two-stage fine-tuning strategy to the model's performance. Further analysis investigates the impact of input data sequence length and image quality on both accuracy and inference time, as well as the interpretability of the enhanced CoT reasoning ability achieved through fine-tuning.

Cómo citar

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Y. L. E. (2026). VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction. https://doi.org/10.26599/COMMTR.2026.9640009

MLA

al, Yancheng Ling et. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction." 2026. https://doi.org/10.26599/COMMTR.2026.9640009.

Chicago

al, Yancheng Ling et. 2026. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction.". https://doi.org/10.26599/COMMTR.2026.9640009.

Harvard

al, Y. L. E. 2026, VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction, Tsinghua University Press, available at: https://doi.org/10.26599/COMMTR.2026.9640009 [Accessed 10 Aug. 2026].

Compartir e imprimir

Guardá la ficha, copiá su enlace permanente o imprimila como PDF.

Exportar referencia

Si usás un gestor bibliográfico, podés exportar el registro en los formatos más comunes.

Detalles del recurso

Información bibliográfica útil para confirmar que se trata del material correcto.

Título
VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction
Autor / colaboradores
Yancheng Ling et al
Editorial
Tsinghua University Press
Año de publicación
2026
ISSN
2772-4247
ISSN
2772-4247
Idioma
Inglés

Materias

Explorá otros recursos relacionados a partir de estas materias.

Copiado