Torna ai risultati
Scheda bibliografica · Consultazione e accesso
Artículo

VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction

Yancheng Ling et al · Tsinghua University Press · 2026

Accesso aperto disponibile
Lettura rapida. Controlla i dati essenziali della risorsa e accedi al contenuto con il pulsante principale. La scheda mostra solo le informazioni necessarie per identificare, citare e aprire l’opera.

Accesso alla risorsa

Apri il contenuto dall’opzione principale o scegli un’altra fonte disponibile.

DOAJ DOAJ Articles
Entrar por DOAJ
Accesso principale

Accesso aperto disponibile

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Apri risorsa

Riepilogo

Descripción general del contenido del recurso.

Pedestrian crossing intention prediction is crucial for autonomous driving. While existing models have achieved high accuracy, their generalization and robustness remain limited, hindering their performance in real-world scenarios. To overcome these limitations, we introduce the LVLMPed-CoT, a large vision language model (LVLM) that incorporates a chain-of-thought (CoT) mechanism to enhance pedestrian crossing intention prediction. It takes multimodal data as input and employs data distillation along with a two stage fine-tuning strategy to elicit the implicit CoT capability of a lightweight vision-language model for enhanced perception, reasoning, and prediction. The unified LVLMPed-CoT is trained on a joint open-source dataset (JAAD and PIE) and achieves superior or comparable performance to state-of-the-art models on both large-scale public datasets. The ablation study validates the contribution of the CoT prompt design and the two-stage fine-tuning strategy to the model's performance. Further analysis investigates the impact of input data sequence length and image quality on both accuracy and inference time, as well as the interpretability of the enhanced CoT reasoning ability achieved through fine-tuning.

Come citare

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Y. L. E. (2026). VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction. https://doi.org/10.26599/COMMTR.2026.9640009

MLA

al, Yancheng Ling et. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction." 2026. https://doi.org/10.26599/COMMTR.2026.9640009.

Chicago

al, Yancheng Ling et. 2026. "VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction.". https://doi.org/10.26599/COMMTR.2026.9640009.

Harvard

al, Y. L. E. 2026, VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction, Tsinghua University Press, available at: https://doi.org/10.26599/COMMTR.2026.9640009 [Accessed 7 Aug. 2026].

Condividi e stampa

Salva la scheda, copia il link permanente o stampala in PDF.

Esporta riferimento

Esporta il record nei formati più comuni per usarlo con un gestore bibliografico.

Dettagli della risorsa

Informazioni bibliografiche utili per verificare che sia il materiale corretto.

Titolo
VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction
Autore / collaboratori
Yancheng Ling et al
Editore
Tsinghua University Press
Anno di pubblicazione
2026
ISSN
2772-4247
ISSN
2772-4247
Lingua
Inglés

Soggetti

Esplora risorse correlate a partire da questi soggetti.

Copiato