activity
20172026
most citedThe 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation

100 citations · 147 across the 8 of their papers we have counts for

collaborators

12 papers

cs.CV2026

Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition

Mauro Comi, Jordi Serrano Berbel, Kevis-Kokitsi Maninis +2

While state-of-the-art generative models produce high-fidelity 3D meshes, these outputs lack the physical properties required for interactive simulation, gaming, or robotics. We in…

cs.CV2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis +16

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation…

cs.RO20251 cited

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan +169

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the G…

cs.CV2024

EgoCast: Forecasting Egocentric Human Pose in the Wild

Maria Escobar, Juanita Puentes, Cristhian Forigua +3

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoC…

cs.CV2024

TIPS: Text-Image Pretraining with Spatial awareness

Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh +11

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense und…

cs.CV2019100 cited

The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation

Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi +3

We present the 2019 DAVIS Challenge on Video Object Segmentation, the third edition of the DAVIS Challenge series, a public competition designed for the task of Video Object Segmen…