Search papers, labs, and topics across Lattice.
This study introduces a sim-to-real framework for tomato plant segmentation that leverages synthetic data generation and fine-tuning of the Segment Anything Model 3 (SAM 3). By creating a large-scale synthetic dataset that captures diverse environmental conditions and plant characteristics, the authors enhance the model's text-conditioned segmentation capabilities specifically for greenhouse crops. Evaluation on real-world datasets reveals that this approach significantly boosts segmentation performance and model confidence, addressing the critical challenge of limited annotated training data in agricultural settings.
Fine-tuning a foundation model with procedurally generated synthetic data can dramatically enhance segmentation accuracy in complex agricultural environments.
Vision-based automation is an excellent candidate for reducing manual labor in greenhouse crop production and phenotyping. However, progress is constrained by the lack of annotated training data. Recent advances in vision-based foundational models have shown promising results in zero-shot generalization to novel domains, but their performance drops in complex agricultural environments. In this work, we present a sim-to-real framework for tomato plant segmentation that combines synthetic data generation with fine-tuning of a foundation model. We model a commercial cherry tomato greenhouse and use it to generate a large-scale synthetic dataset under diverse viewpoints, lighting conditions, and plant morphology. Subsequently, we fine-tune the Segment Anything Model 3 (SAM 3) on the synthetic dataset, specializing its text-conditioned segmentation behavior for greenhouse crop organs while retaining the general visual prior that makes zero-shot transfer possible. By evaluating our framework on multiple real-world greenhouse datasets, we demonstrate that combining synthetic data with SAM 3 fine-tuning significantly improves segmentation performance and model confidence. To support community benchmarking, we publicly release the procedural model, the generated synthetic dataset, and our fine-tuned SAM 3 weights.