From the 1 of 13 linked papers with an AI index.
13 papers
Autoregressive Modeling of Film with Applications in Video Montage
Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar +5
FilmGPT is an autoregressive transformer trained on a large movie corpus to learn the statistical patterns of film editing and select existing raw shots to create coherent video mo…
AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
Anna Šárová MikeÅ¡tÃková, Médéric Fourmy, Martin CÃfka +2
Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter, and occlusions. Multi-view pose…
Temporally Consistent Object 6D Pose Estimation for Robot Control
Kateryna Zorina, Vojtech Priban, Mederic Fourmy +2
Single-view RGB object pose estimators have reached a level of precision and efficiency that makes them good candidates for vision-based robot control. However, off-the-shelf metho…
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
Jai Bardhan, Patrik Drozdik, Josef Sivic +1
Action-conditioned robot world models generate future video frames of the manipulated scene given a robot action sequence, offering a promising alternative for simulating tasks tha…
REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
Martin Sedlacek, Pavlo Yefanov, Georgy Ponimatkin +7
Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to gen…
ResidualViT for Efficient Temporally Dense Video Encoding
Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2
Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…