Amodal instance segmentation predicts complete object shapes, including hidden regions, which is crucial for accurate fruit sizing and counting in the agri-food Industry 4.0. However, amodal models require massive datasets of labor-intensive annotations, hindering deployment across new orchards. This study investigates if geometric rules of occlusion reasoning can transfer from densely annotated urban driving scenes to agricultural environments. We fine-tune AISFormer, a transformer-based amodal model pre-trained on the KINS2020 benchmark, over the PApple orchard dataset. Experiments reveal cross-domain knowledge transfer is remarkably data-efficient despite the extreme visual domain gap. Using the full dataset, our approach achieves a state-of-the-art F1 score of 0.875. Crucially, we match the ORCNN baseline (F1=0.860) using only 43.4% of the requisite training data, and recover 94.3% of ORCNN baseline performance from 100 annotated images. These findings provide a sustainable, high-efficiency pathway for deploying precise amodal vision systems while drastically reducing annotation costs.
AI-Driven Amodal Instance Segmentation for Precision Agriculture: Data-Efficient Apple Detection via Cross-Domain Transfer Learning
Tarif, Mehran;Saadatpour, Mohsen;Quaglia, Davide
2026-01-01
Abstract
Amodal instance segmentation predicts complete object shapes, including hidden regions, which is crucial for accurate fruit sizing and counting in the agri-food Industry 4.0. However, amodal models require massive datasets of labor-intensive annotations, hindering deployment across new orchards. This study investigates if geometric rules of occlusion reasoning can transfer from densely annotated urban driving scenes to agricultural environments. We fine-tune AISFormer, a transformer-based amodal model pre-trained on the KINS2020 benchmark, over the PApple orchard dataset. Experiments reveal cross-domain knowledge transfer is remarkably data-efficient despite the extreme visual domain gap. Using the full dataset, our approach achieves a state-of-the-art F1 score of 0.875. Crucially, we match the ORCNN baseline (F1=0.860) using only 43.4% of the requisite training data, and recover 94.3% of ORCNN baseline performance from 100 annotated images. These findings provide a sustainable, high-efficiency pathway for deploying precise amodal vision systems while drastically reducing annotation costs.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



