Back to results
Bibliographic record · Consultation and access
Artículo

MoTIF: An end-to-end multimodal road traffic scene understanding foundation model

Zihe Wang et al · Tsinghua University Press · 2025

Open access available
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

DOAJ DOAJ Articles
Entrar por DOAJ
Main access

Open access available

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Open resource

Summary

Descripción general del contenido del recurso.

Video-based road intelligent detection constitutes a critical component in modern intelligent transportation systems, serving as a crucial role for comprehensive transportation planning and emergency traffic management. Current traffic scene perception methodologies relying on conventional deep learning architectures present inherent limitations, including heavy dependence on extensive manual annotations of specific traffic scenarios and predefined rule configurations. These approaches demonstrate constrained semantic representation capacity and limited generalizability across heterogeneous traffic scenarios. To address these challenges, this study proposes a novel end-to-end multimodal foundation model architecture that jointly generates dynamic traffic event detection outcomes and semantic-rich contextual descriptions. Through integration of low-rank adaptation (LoRA) and prompt fine-tuning as parameter-efficient fine-tuning strategies, we develop the multimodal road traffic scene understanding foundation model (MoTIF), which establishes cross-modal alignment between visual patterns and textual semantics. This framework demonstrates enhanced capability in extracting salient traffic targets and generating hierarchical scene representations, significantly improving automated detection efficiency in road video analytics. Notably, MoTIF exhibits contextual reasoning capabilities for implicit traffic event interpretation. Extensive evaluations on two real-world datasets encompassing urban road intersection scenarios in Tianjin and highway monitoring systems in Shandong Province reveal that MoTIF achieves superior performance metrics: 65.81 average score on multimodal scene understanding assessment and 83.33% event detection accuracy, outperforming mainstream benchmarks in both precision and computational efficiency. This research advances multimodal learning paradigms for intelligent transportation systems while providing practical insights for adaptive traffic management applications.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, Z. W. E. (2025). MoTIF: An end-to-end multimodal road traffic scene understanding foundation model. https://doi.org/10.1016/j.commtr.2025.100227

MLA

al, Zihe Wang et. "MoTIF: An end-to-end multimodal road traffic scene understanding foundation model." 2025. https://doi.org/10.1016/j.commtr.2025.100227.

Chicago

al, Zihe Wang et. 2025. "MoTIF: An end-to-end multimodal road traffic scene understanding foundation model.". https://doi.org/10.1016/j.commtr.2025.100227.

Harvard

al, Z. W. E. 2025, MoTIF: An end-to-end multimodal road traffic scene understanding foundation model, Tsinghua University Press, available at: https://doi.org/10.1016/j.commtr.2025.100227 [Accessed 7 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
MoTIF: An end-to-end multimodal road traffic scene understanding foundation model
Author / contributors
Zihe Wang et al
Publisher
Tsinghua University Press
Publication year
2025
ISSN
2772-4247
ISSN
2772-4247
Language
English

Subjects

Explore related resources through these subjects.

Copied