3 citations · 3 across the 7 of their papers we have counts for
9 papers
Temporally Consistent Object 6D Pose Estimation for Robot Control
Kateryna Zorina, Vojtech Priban, Mederic Fourmy +2
Single-view RGB object pose estimators have reached a level of precision and efficiency that makes them good candidates for vision-based robot control. However, off-the-shelf metho…
REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
Martin Sedlacek, Pavlo Yefanov, Georgy Ponimatkin +7
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to gen…
ResidualViT for Efficient Temporally Dense Video Encoding
Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2
Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…
Discovering Divergent Representations between Text-to-Image Models
Lisa Dunlap, Joseph E. Gonzalez, Trevor Darrell +3
In this paper, we investigate when and how visual representations learned by two different generative models diverge. Given two text-to-image models, our goal is to discover visual…
Improving Personalized Search with Regularized Low-Rank Parameter Updates
Fiona Ryan, Josef Sivic, Fabian Caba Heilbron +3
Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning…
Multi-step manipulation task and motion planning guided by video demonstration
Kateryna Zorina, David Kovar, Mederic Fourmy +5
This work aims to leverage instructional video to solve complex multi-step task-and-motion planning tasks in robotics. Towards this goal, we propose an extension of the well-establ…