Volver a resultados
Ficha bibliográfica · Consulta y acceso
Artículo de revista

Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing

Xing Zi et al · IEEE · 2026

Acceso abierto disponible
Lectura rápida. Revisá los datos básicos del recurso y luego accedé al contenido desde el botón principal. En esta ficha solo se muestra la información necesaria para identificar la obra, citarla y abrirla.
Publicación seriada

3PS-RAN: A Real-Time Framework for Securing the O-RAN RACH Against DDoS Attacks Toward NextG

Esta publicación seriada contiene 172 contenidos relacionados.

Acceso al recurso

Entrá al contenido desde la opción principal o elegí otra fuente disponible.

DOAJ DOAJ Articles
Entrar por DOAJ
Acceso principal

Acceso abierto disponible

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Abrir recurso

Resumen

Descripción general del contenido del recurso.

Pixel-level segmentation is critical for remote sensing applications, yet traditional supervised methods suffer from high annotation costs. While foundational vision models like CLIP and the Segment Anything Model (SAM) offer zero-shot capabilities, they struggle with domain-specific challenges in aerial imagery, specifically: (1) scale variation causing attention drift, (2) lack of semantic discrimination leading to mask redundancy, and (3) poor adaptation to overhead perspectives. To bridge this gap without task-specific fine-tuning, VTPSeg is presented as a coarse-to-fine multi-model framework designed for high-precision off-line mapping. Unlike generic integrations, VTPSeg introduces a cohesive semantic-geometric synergy. Specifically, the Grounding DINO+ (GD+) module employs a novel synonym-based prompt strategy to maximize recall for overhead objects. The CLIP Filter++ module then utilizes a dual-prompt mechanism (visual attention circles and negative text constraints) to eliminate false positives caused by background clutter. Finally, these refined priors serve as precise point prompts for FastSAM, ensuring instance-level granularity. Validated on five diverse datasets (WHU, LoveDA, Inria, xBD, and iSAID), VTPSeg achieves state-of-the-art or highly competitive performance across five diverse datasets, demonstrating that strategic prompt engineering can effectively adapt frozen foundational models to complex remote sensing tasks.

Cómo citar

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, X. Z. E. (2026). Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing. https://doi.org/10.1109/ACCESS.2026.3684906

MLA

al, Xing Zi et. "Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing." 2026. https://doi.org/10.1109/ACCESS.2026.3684906.

Chicago

al, Xing Zi et. 2026. "Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing.". https://doi.org/10.1109/ACCESS.2026.3684906.

Harvard

al, X. Z. E. 2026, Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing, IEEE, available at: https://doi.org/10.1109/ACCESS.2026.3684906 [Accessed 10 Aug. 2026].

Compartir e imprimir

Guardá la ficha, copiá su enlace permanente o imprimila como PDF.

Exportar referencia

Si usás un gestor bibliográfico, podés exportar el registro en los formatos más comunes.

Detalles del recurso

Información bibliográfica útil para confirmar que se trata del material correcto.

Título
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
Autor / colaboradores
Xing Zi et al
Editorial
IEEE
Año de publicación
2026
ISSN
2169-3536
ISSN
2169-3536
Idioma
Inglés

Materias

Explorá otros recursos relacionados a partir de estas materias.

Copiado