Zurück zu den Ergebnissen
Bibliografischer Datensatz · Ansicht und Zugriff
Artículo de revista

Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation

Ruize Xia · IEEE · 2026

Ergänzendes Material verfügbar
Schnellübersicht. Prüfen Sie die grundlegenden Angaben und öffnen Sie den Inhalt über die Hauptschaltfläche. Die Seite zeigt nur die Informationen, die zum Identifizieren, Zitieren und Öffnen des Werks nötig sind.
Fortlaufende Publikation

3PS-RAN: A Real-Time Framework for Securing the O-RAN RACH Against DDoS Attacks Toward NextG

Diese fortlaufende Publikation enthält 172 zugehörige Inhalte.

Zugriff auf die Ressource

Öffnen Sie den Inhalt über die Hauptoption oder wählen Sie eine andere verfügbare Quelle.

DOAJ DOAJ Articles
Entrar por DOAJ
Hauptzugriff

Ergänzendes Material verfügbar

El enlace apunta a material asociado, anexos, tablas, datos o página complementaria. No se marca como libro/texto completo.
Material öffnen

Übersicht

Descripción general del contenido del recurso.

Sign language is a primary communication channel for millions of people who are deaf or hard of hearing, yet generating signer video directly from text remains difficult because video diffusion models are expensive to train and evaluate. This paper presents Text2Sign, a text-conditioned diffusion architecture for short sign-language video clips, designed to operate on a single NVIDIA L4 graphics processor rather than on a multi-node training infrastructure. The model combines a frozen vision&#x2013;language text encoder with a three-dimensional encoder&#x2013;decoder backbone and factorized spatial and temporal attention, thereby reducing the cost of full spatio-temporal attention while preserving motion coherence. Three design choices are examined: whether transformer-style blocks improve upon convolution-only baselines, whether a frozen pretrained text encoder yields lower loss than a task-specific encoder trained from scratch under the present short-budget comparison, and whether factorized attention is competitive with full video attention. On a signer-disjoint partition of short clips extracted from How2Sign, the best short-run ablation attains a validation loss of 0.0648, while a longer-run checkpoint reaches 0.00999. A compact evaluation slice of that checkpoint yields SSIM <inline-formula> <tex-math notation="LaTeX">$0.2403\pm 0.0238$ </tex-math></inline-formula>, PSNR <inline-formula> <tex-math notation="LaTeX">$15.11\pm 0.42$ </tex-math></inline-formula>&#x2006;dB, and temporal consistency <inline-formula> <tex-math notation="LaTeX">$1.0000\pm 0.0000$ </tex-math></inline-formula>; under an 8-step DDIM setting with guidance scale 5.0, the model generates a 32-frame <inline-formula> <tex-math notation="LaTeX">$64\times 64$ </tex-math></inline-formula> clip in 12.60&#x2006;s (2.54 frames/s) with 3.12&#x2006;GB peak inference memory on a single NVIDIA L4. In a held-out conditional denoising audit on real validation clips, removing text raises late-timestep denoising loss from 0.9875 to 0.9891, whereas shuffled prompts remain nearly indistinguishable from the intended prompt. Thus, frozen text conditioning yields a lower short-budget validation loss than the custom encoder baseline, and the revised post-revision checkpoint is qualitatively stronger than the earlier baseline in direct side-by-side inspection; however, held-out audits still show only weak prompt-specific separation. The present system remains limited to low-resolution short clips and does not yet include expert linguistic evaluation; accordingly, the reported results should be interpreted as a single-GPU research baseline rather than a complete solution to sign-language production. The code is publicly available at <uri>https://github.com/xiaruize0911/text2sign</uri>

Zitieren

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

Xia, R. (2026). Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation. https://doi.org/10.1109/ACCESS.2026.3686260

MLA

Xia, Ruize. "Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation." 2026. https://doi.org/10.1109/ACCESS.2026.3686260.

Chicago

Xia, Ruize. 2026. "Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation.". https://doi.org/10.1109/ACCESS.2026.3686260.

Harvard

Xia, R. 2026, Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation, IEEE, available at: https://doi.org/10.1109/ACCESS.2026.3686260 [Accessed 7 Aug. 2026].

Teilen und drucken

Speichern Sie den Datensatz, kopieren Sie den Permalink oder drucken Sie ihn als PDF.

Referenz exportieren

Exportieren Sie den Datensatz in gängigen Formaten für Literaturverwaltungsprogramme.

Ressourcendetails

Bibliografische Angaben zur Prüfung, ob es sich um das richtige Material handelt.

Titel
Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
Autor / Mitwirkende
Ruize Xia
Verlag
IEEE
Erscheinungsjahr
2026
ISSN
2169-3536
ISSN
2169-3536
Sprache
Inglés

Schlagwörter

Entdecken Sie über diese Schlagwörter weitere verwandte Ressourcen.

Kopiert