44 citations · 76 across the 8 of their papers we have counts for
8 papers
Revealing data leakage in protein interaction benchmarks
Anton Bushuiev, Roman Bushuiev, Jiri Sedlar +4
In recent years, there has been remarkable progress in machine learning for protein-protein interactions. However, prior work has predominantly focused on improving learning algori…
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
Antonin Vobecky, Oriane Siméoni, David Hurych +4
We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of f…
Visually Guided Model Predictive Robot Control via 6D Object Pose Localization and Tracking
Mederic Fourmy, Vojtech Priban, Jan Kristof Behrens +3
The objective of this work is to enable manipulation tasks with respect to the 6D pose of a dynamically moving object using a camera mounted on a robot. Examples include maintainin…
VidChapters-7M: Video Chapters at Scale
Antoine Yang, Arsha Nagrani, Ivan Laptev +2
Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly…
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5
In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architec…
Learning Object Manipulation Skills from Video via Approximate Differentiable Physics
Vladimir Petrik, Mohammad Nomaan Qureshi, Josef Sivic +1
We aim to teach robots to perform simple object manipulation tasks by watching a single video demonstration. Towards this goal, we propose an optimization approach that outputs a c…