Torna ai risultati
Scheda bibliografica · Consultazione e accesso
Artículo

Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images

Olivia Zumsteg et al · Elsevier · 2026

Accesso aperto disponibile
Lettura rapida. Controlla i dati essenziali della risorsa e accedi al contenuto con il pulsante principale. La scheda mostra solo le informazioni necessarie per identificare, citare e aprire l’opera.

Accesso alla risorsa

Apri il contenuto dall’opzione principale o scegli un’altra fonte disponibile.

DOAJ DOAJ Articles
Entrar por DOAJ
Accesso principale

Accesso aperto disponibile

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Apri risorsa

Riepilogo

Descripción general del contenido del recurso.

Estimating three-dimensional morphological traits such as volume from two-dimensional RGB images presents inherent challenges due to the loss of depth information, projection distortions, and occlusions under field conditions. In this work, we explore multiple approaches for non-destructive volume estimation of wheat spikes using RGB images and structured-light 3D scans as ground truth references. Wheat spike volume is promising for phenotyping as it shows a high correlation with spike dry weight, a key component of fruiting efficiency. We define the total spike volume per unit area of a wheat canopy at flowering as fruiting capacity. Accounting for the complex geometry of the spikes, we compare different neural network approaches for volume estimation from 2D images and benchmark them against two conventional baselines: a 2D area-based projection and a geometric reconstruction using axis-aligned cross-sections. Fine-tuned Vision Transformers (DINOv2 and DINOv3) with MLPs achieve the lowest MAPE of 5.08% and 4.67%, and the highest correlation of 0.96 and 0.97 on six-view indoor images, outperforming fine-tuned CNNs (ResNet18 and ResNet50), a wheat-specific backbone, and both baselines. When using frozen DINO backbones, deep-supervised LSTMs outperform MLPs, whereas after fine-tuning, improved high-level representations allow simple MLPs to outperform LSTMs. We demonstrate that object shape significantly impacts volume estimation accuracy, with irregular geometries such as wheat spikes posing greater challenges for geometric methods than for deep learning approaches. Fine-tuning DINOv3 on field-based single side-view images yields a MAPE of 8.39% and a correlation of 0.90, providing a novel pipeline and a fast, accurate, and non-destructive approach for wheat spike volume phenotyping.

Come citare

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, O. Z. E. (2026). Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images. https://doi.org/10.1016/j.atech.2026.102150

MLA

al, Olivia Zumsteg et. "Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images." 2026. https://doi.org/10.1016/j.atech.2026.102150.

Chicago

al, Olivia Zumsteg et. 2026. "Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images.". https://doi.org/10.1016/j.atech.2026.102150.

Harvard

al, O. Z. E. 2026, Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images, Elsevier, available at: https://doi.org/10.1016/j.atech.2026.102150 [Accessed 8 Aug. 2026].

Condividi e stampa

Salva la scheda, copia il link permanente o stampala in PDF.

Esporta riferimento

Esporta il record nei formati più comuni per usarlo con un gestore bibliografico.

Dettagli della risorsa

Informazioni bibliografiche utili per verificare che sia il materiale corretto.

Titolo
Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images
Autore / collaboratori
Olivia Zumsteg et al
Editore
Elsevier
Anno di pubblicazione
2026
ISSN
2772-3755
ISSN
2772-3755
Lingua
Inglés

Soggetti

Esplora risorse correlate a partire da questi soggetti.

Copiato