Torna ai risultati
Scheda bibliografica · Consultazione e accesso
Artículo

MoTIF: An end-to-end multimodal road traffic scene understanding foundation model

Zihe Wang et al · Tsinghua University Press · 2025

Accesso aperto disponibile
Lettura rapida. Controlla i dati essenziali della risorsa e accedi al contenuto con il pulsante principale. La scheda mostra solo le informazioni necessarie per identificare, citare e aprire l’opera.

Accesso alla risorsa

Apri il contenuto dall’opzione principale o scegli un’altra fonte disponibile.

DOAJ DOAJ Articles
Entrar por DOAJ
Accesso principale

Accesso aperto disponibile

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Apri risorsa

Riepilogo

Descripción general del contenido del recurso.

Video-based road intelligent detection constitutes a critical component in modern intelligent transportation systems, serving as a crucial role for comprehensive transportation planning and emergency traffic management. Current traffic scene perception methodologies relying on conventional deep learning architectures present inherent limitations, including heavy dependence on extensive manual annotations of specific traffic scenarios and predefined rule configurations. These approaches demonstrate constrained semantic representation capacity and limited generalizability across heterogeneous traffic scenarios. To address these challenges, this study proposes a novel end-to-end multimodal foundation model architecture that jointly generates dynamic traffic event detection outcomes and semantic-rich contextual descriptions. Through integration of low-rank adaptation (LoRA) and prompt fine-tuning as parameter-efficient fine-tuning strategies, we develop the multimodal road traffic scene understanding foundation model (MoTIF), which establishes cross-modal alignment between visual patterns and textual semantics. This framework demonstrates enhanced capability in extracting salient traffic targets and generating hierarchical scene representations, significantly improving automated detection efficiency in road video analytics. Notably, MoTIF exhibits contextual reasoning capabilities for implicit traffic event interpretation. Extensive evaluations on two real-world datasets encompassing urban road intersection scenarios in Tianjin and highway monitoring systems in Shandong Province reveal that MoTIF achieves superior performance metrics: 65.81 average score on multimodal scene understanding assessment and 83.33% event detection accuracy, outperforming mainstream benchmarks in both precision and computational efficiency. This research advances multimodal learning paradigms for intelligent transportation systems while providing practical insights for adaptive traffic management applications.

Come citare

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Z. W. E. (2025). MoTIF: An end-to-end multimodal road traffic scene understanding foundation model. https://doi.org/10.1016/j.commtr.2025.100227

MLA

al, Zihe Wang et. "MoTIF: An end-to-end multimodal road traffic scene understanding foundation model." 2025. https://doi.org/10.1016/j.commtr.2025.100227.

Chicago

al, Zihe Wang et. 2025. "MoTIF: An end-to-end multimodal road traffic scene understanding foundation model.". https://doi.org/10.1016/j.commtr.2025.100227.

Harvard

al, Z. W. E. 2025, MoTIF: An end-to-end multimodal road traffic scene understanding foundation model, Tsinghua University Press, available at: https://doi.org/10.1016/j.commtr.2025.100227 [Accessed 8 Aug. 2026].

Condividi e stampa

Salva la scheda, copia il link permanente o stampala in PDF.

Esporta riferimento

Esporta il record nei formati più comuni per usarlo con un gestore bibliografico.

Dettagli della risorsa

Informazioni bibliografiche utili per verificare che sia il materiale corretto.

Titolo
MoTIF: An end-to-end multimodal road traffic scene understanding foundation model
Autore / collaboratori
Zihe Wang et al
Editore
Tsinghua University Press
Anno di pubblicazione
2025
ISSN
2772-4247
ISSN
2772-4247
Lingua
Inglés

Soggetti

Esplora risorse correlate a partire da questi soggetti.

Copiato