From the 2 of 5 linked papers with an AI index.
5 papers
ContactFlow: A video action conditioning that transfers across embodiments
Sami Azirar, Enrico Pallotta, Jan Nogga +3
The paper introduces Contact Flow, an embodiment‑agnostic representation that encodes manipulation as the trajectory of 3D contact points, enabling a video‑based world model traine…
Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data
Enrico Pallotta, Mohamed Farag, Esra Guclu +3
Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics…
Still image and spatial-temporal tomato data enabling detection, segmentation, tracking, and video-instance segmentation using strong and weak labels
Michael Halstead, Esra Guclu, Mohamed Farag +7
The paper introduces two new datasets of tomato plants captured by a robot—still images (BUTom21) and video sequences (BUTom-ST21)—with pixel‑level annotations for fruit detection,…
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta +3
In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuo…
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos +3
Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work…