Back to results
Bibliographic record · Consultation and access
Artículo

VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction

Yancheng Ling et al · Tsinghua University Press · 2026

Open access available
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

DOAJ DOAJ Articles
Entrar por DOAJ
Main access

Open access available

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Open resource

Summary

Descripción general del contenido del recurso.

Pedestrian crossing intention prediction is crucial for autonomous driving. While existing models have achieved high accuracy, their generalization and robustness remain limited, hindering their performance in real-world scenarios. To overcome these limitations, we introduce the LVLMPed-CoT, a large vision language model (LVLM) that incorporates a chain-of-thought (CoT) mechanism to enhance pedestrian crossing intention prediction. It takes multimodal data as input and employs data distillation along with a two stage fine-tuning strategy to elicit the implicit CoT capability of a lightweight vision-language model for enhanced perception, reasoning, and prediction. The unified LVLMPed-CoT is trained on a joint open-source dataset (JAAD and PIE) and achieves superior or comparable performance to state-of-the-art models on both large-scale public datasets. The ablation study validates the contribution of the CoT prompt design and the two-stage fine-tuning strategy to the model's performance. Further analysis investigates the impact of input data sequence length and image quality on both accuracy and inference time, as well as the interpretability of the enhanced CoT reasoning ability achieved through fine-tuning.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Y. L. E. (2026). VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction. https://doi.org/10.26599/COMMTR.2026.9640009

MLA

al, Yancheng Ling et. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction." 2026. https://doi.org/10.26599/COMMTR.2026.9640009.

Chicago

al, Yancheng Ling et. 2026. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction.". https://doi.org/10.26599/COMMTR.2026.9640009.

Harvard

al, Y. L. E. 2026, VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction, Tsinghua University Press, available at: https://doi.org/10.26599/COMMTR.2026.9640009 [Accessed 7 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction
Author / contributors
Yancheng Ling et al
Publisher
Tsinghua University Press
Publication year
2026
ISSN
2772-4247
ISSN
2772-4247
Language
English

Subjects

Explore related resources through these subjects.

Copied