AI summary

75% confidence

The proposed MSA-MVSNet network addresses issues in 3D reconstruction of orchard trees by utilizing a cross-scale collaborative attention mechanism and integrating semantic segmentation for fruit counting. It adapts to complex branch-leaf morphology and illumination variations, enhancing detail representation. The network improves modeling robustness and matching stability.

Generated by MESSAI extraction pipeline · review against source PDF

Literature priors

Distribution

Reported parameters

The author’s reported value (▼) sits on top of the literature distribution from MESS-Parameters. Values outside the band are flagged as outliers.

Coulombic efficiency3.5%

2 extracted values

View extracted values

No 3D model is mapped to this paper yet. Parameter ranges above still place reported values on the literature distribution.

Abstract

Abstract To address the issues of detail loss and matching difficulties in fruit tree 3D reconstruction caused by complex branch–leaf morphology, fruit occlusion, and illumination variations, this paper proposes an end-to-end cross-scale collaborative attention multi-view stereo network, termed MSA-MVSNet, for high-quality 3D reconstruction of orchard trees, while integrating semantic segmentation for fruit counting. A multi-scale feature enhancement module is designed to adaptively fuse deep semantic features and shallow fine-grained details through a spatial–channel collaborative attention mechanism, thereby enhancing the network’s capability to represent multi-scale structures such as trunks, branches, and leaves. Multi-branch dilated convolutions are introduced to enlarge the receptive field, and deformable convolutions are incorporated to adaptively capture the irregular geometric shapes of fruits, improving modeling robustness. In addition, a feature matching transformer is introduced to strengthen long-range global contextual correlations within and across images via intra-attention and inter-attention mechanisms, thereby improving matching stability in low-texture and repetitive-texture regions.To validate the effectiveness of the proposed method, experiments are conducted on self-collected real orchard dataset and public benchmark datasets. The results demonstrate that MSA-MVSNet outperforms baseline models by 8.2% in terms of 3D reconstruction quality. Finally, by combining depth filtering with the semantic segmentation results of YOLOv11-Seg, a semantic-guided fruit reconstruction and counting framework is constructed. This framework achieves an overall counting F1-score of 92.8% on the self-collected dataset with varying scene sparsity and 93.5% on the public Fuji-sfm dataset, demonstrating its effectiveness and generalization capability.

Key findings

  • MSA-MVSNet achieves high-quality 3D reconstruction of orchard trees
  • The network effectively integrates semantic segmentation for fruit counting
  • The spatial-channel collaborative attention mechanism enhances multi-scale structure representation

Keywords

Artificial intelligenceSegmentationBenchmark (surveying)3D reconstructionComputer visionFeature (linguistics)

Identifiers

Journal
Research Square
Year
2026