11 papers
BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
Jakub Zadrozny, Oisin Mac Aodha, Hakan Bilen
3D reconstruction of articulated objects from a single image is challenging because large training datasets with paired image and 3D supervision are difficult to obtain. Recent poi…
BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
Miaowei Wang, Qingxuan Yan, Zhi Cao +4
Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing…
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +4
The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. Howev…
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
Duolikun Danier, Ge Gao, Steven McDonagh +3
Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated vid…
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +2
The integration of pre-trained visual representations (PVRs) has significantly advanced visuomotor policy learning. However, effectively leveraging these models remains a challenge…
Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
Octave Mariotti, Zhipeng Du, Yash Bhalgat +2
Semantic correspondence (SC) aims to establish semantically meaningful matches across different instances of an object category. We illustrate how recent supervised SC methods rema…