11 papers
Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
Philippe Weinzaepfel, Christian Wolf, Bülent Mert Sariyildiz +2
Transformers are AI's workhorse with strong performance in modeling sequential data, but their computational cost becomes prohibitive when processing long sequences. We target long…
Multi-HMR 2: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
Guénolé Fiche, Philippe Weinzaepfel, Romain Brégier +1
Most advances in human mesh recovery (HMR) have focused on pelvis-centered recovery, overlooking metric 3D localization and detection accuracy in the camera coordinate system - two…
Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks
Pau de Jorge, César Roberto de Souza, Björn Michele +5
Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, m…
What does really matter in image goal navigation?
Gianluca Monaci, Philippe Weinzaepfel, Christian Wolf
Image goal navigation requires two different skills: firstly, core navigation skills, including the detection of free space and obstacles, and taking decisions based on an internal…
Human Mesh Modeling for Anny Body
Romain Brégier, Guénolé Fiche, Laura Bravo-Sánchez +5
Parametric body models provide the structural basis for many human-centric tasks, yet existing models often rely on costly 3D scans and learned shape spaces that are proprietary an…
Kinaema: a recurrent sequence model for memory and pose in motion
Mert Bulent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono +2
One key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves in previously seen spaces. In this work, we focus on this part…