Back to results
Bibliographic record · Consultation and access
Artículo

Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images

Olivia Zumsteg et al · Elsevier · 2026

Open access available
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

DOAJ DOAJ Articles
Entrar por DOAJ
Main access

Open access available

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Open resource

Summary

Descripción general del contenido del recurso.

Estimating three-dimensional morphological traits such as volume from two-dimensional RGB images presents inherent challenges due to the loss of depth information, projection distortions, and occlusions under field conditions. In this work, we explore multiple approaches for non-destructive volume estimation of wheat spikes using RGB images and structured-light 3D scans as ground truth references. Wheat spike volume is promising for phenotyping as it shows a high correlation with spike dry weight, a key component of fruiting efficiency. We define the total spike volume per unit area of a wheat canopy at flowering as fruiting capacity. Accounting for the complex geometry of the spikes, we compare different neural network approaches for volume estimation from 2D images and benchmark them against two conventional baselines: a 2D area-based projection and a geometric reconstruction using axis-aligned cross-sections. Fine-tuned Vision Transformers (DINOv2 and DINOv3) with MLPs achieve the lowest MAPE of 5.08% and 4.67%, and the highest correlation of 0.96 and 0.97 on six-view indoor images, outperforming fine-tuned CNNs (ResNet18 and ResNet50), a wheat-specific backbone, and both baselines. When using frozen DINO backbones, deep-supervised LSTMs outperform MLPs, whereas after fine-tuning, improved high-level representations allow simple MLPs to outperform LSTMs. We demonstrate that object shape significantly impacts volume estimation accuracy, with irregular geometries such as wheat spikes posing greater challenges for geometric methods than for deep learning approaches. Fine-tuning DINOv3 on field-based single side-view images yields a MAPE of 8.39% and a correlation of 0.90, providing a novel pipeline and a fast, accurate, and non-destructive approach for wheat spike volume phenotyping.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, O. Z. E. (2026). Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images. https://doi.org/10.1016/j.atech.2026.102150

MLA

al, Olivia Zumsteg et. "Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images." 2026. https://doi.org/10.1016/j.atech.2026.102150.

Chicago

al, Olivia Zumsteg et. 2026. "Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images.". https://doi.org/10.1016/j.atech.2026.102150.

Harvard

al, O. Z. E. 2026, Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images, Elsevier, available at: https://doi.org/10.1016/j.atech.2026.102150 [Accessed 6 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
Fine-tuned vision transformers capture complex wheat spike morphology for volume estimation from RGB images
Author / contributors
Olivia Zumsteg et al
Publisher
Elsevier
Publication year
2026
ISSN
2772-3755
ISSN
2772-3755
Language
English

Subjects

Explore related resources through these subjects.

Copied