computer vision

Still image and spatial-temporal tomato data enabling detection, segmentation, tracking, and video-instance segmentation using strong and weak labels

arXiv:2607.14934

summary

The paper introduces two new datasets of tomato plants captured by a robot—still images (BUTom21) and video sequences (BUTom-ST21)—with pixel‑level annotations for fruit detection, segmentation, tracking, and ripeness classification.

Abstract

In this manuscript we release two datasets for visual sensing of tomato plants grown in commercial-like settings and acquired using a robot. The first is BUTom21 which consists of still images and manual annotations. The second is BUTom-ST21 which consists of video-based data and semi-automated annotations through AI-based methods, referred to as pseudo-labels. In both cases, we provide pixel-level labels for the ripeness of the fruit. The aim is to provide the research community a challenging set of real-world imagery to explore methods to sense and estimate the state of tomato plants and their fruit, which is an important horticultural crop. Importantly, the spatial-temporal dataset provides individual fruit count and ripeness information enabling researchers to push the boundaries of field-based phenotyping.

21 pages, 2 figures, 9 tables. Two novel datasets released - link to repository in document

Topics & keywords

#tomato phenotyping#image dataset#video dataset#semantic segmentation#object detection#trackingpixel-level annotationsripeness labelingpseudo-labelsspatial-temporal datarobotic acquisitionfield phenotyping