Accurate 3D-model-based plant phenotyping requires high-quality images with precise overlap, which are difficult to obtain by non-professional people in a crop field. We propose VISAR, a lightweight video-based frame selection pipeline that enables non-expert users to generate accurate 3D reconstructions using casual smartphone video recordings. The pipeline includes five stages: (1) frame extraction, (2) quality filtering via BRISQUE, (3) viewpoint diversity through feature comparison, (4) continuity recovery for dissimilar frames, and (5) plant segmentation for data augmentation. Redundant frames (80% or more similarity) are discarded, frames with moderate similarity (70%-80%) are retained, and bridging frames are added when similarity drops below 70%. Experiments on pepper and cabbage plant datasets demonstrate robust cross-species performance with 2.3-2.7 times higher efficiency than basic video processing. Cabbage achieves 2.7x improvement (38.2 vs 14.2 points/frame) while pepper shows 2.3x improvement (21.1 vs 9.1 points/frame). Both species demonstrate significant quality improvements: pepper achieves 69% hole reduction (16 to 5) and cabbage shows 33% hole reduction (12 to 8), while both achieve superior plant completeness with the whole pipeline (pepper: 83.6%, cabbage: 89.1%). The proposed approach bridges expert photogrammetry and accessible agricultural use, supporting effective plant monitoring with minimal effort.
VISAR: Intelligent Video Frame Selection for Agricultural 3D Plant Reconstruction
Tarif, Mehran;Fasani, Elisa;Quaglia, Davide
2025-01-01
Abstract
Accurate 3D-model-based plant phenotyping requires high-quality images with precise overlap, which are difficult to obtain by non-professional people in a crop field. We propose VISAR, a lightweight video-based frame selection pipeline that enables non-expert users to generate accurate 3D reconstructions using casual smartphone video recordings. The pipeline includes five stages: (1) frame extraction, (2) quality filtering via BRISQUE, (3) viewpoint diversity through feature comparison, (4) continuity recovery for dissimilar frames, and (5) plant segmentation for data augmentation. Redundant frames (80% or more similarity) are discarded, frames with moderate similarity (70%-80%) are retained, and bridging frames are added when similarity drops below 70%. Experiments on pepper and cabbage plant datasets demonstrate robust cross-species performance with 2.3-2.7 times higher efficiency than basic video processing. Cabbage achieves 2.7x improvement (38.2 vs 14.2 points/frame) while pepper shows 2.3x improvement (21.1 vs 9.1 points/frame). Both species demonstrate significant quality improvements: pepper achieves 69% hole reduction (16 to 5) and cabbage shows 33% hole reduction (12 to 8), while both achieve superior plant completeness with the whole pipeline (pepper: 83.6%, cabbage: 89.1%). The proposed approach bridges expert photogrammetry and accessible agricultural use, supporting effective plant monitoring with minimal effort.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



