Publications (54)
3D-CODED : 3D Correspondences by Deep Deformation
Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2
We present a new deep learning approach for matching deformable shapes by introducing {\it Shape Deformation Networks} which jointly encode 3D shapes and correspondences. This is a…
General Detection-based Text Line Recognition
Raphael Baena, Syrine Kalleli, Mathieu Aubry
We introduce a general detection-based approach to text line recognition, be it printed (OCR) or handwritten (HTR), with Latin, Chinese, or ciphered characters. Detection-based app…
AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation
Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2
We introduce a method for learning to generate the surface of 3D shapes. Our approach represents a 3D shape as a collection of parametric surface elements and, in contrast to metho…
Satellite Image Time Series Semantic Change Detection: Novel Architecture and Analysis of Domain Shift
Elliot Vincent, Jean Ponce, Mathieu Aubry
Satellite imagery plays a crucial role in monitoring changes happening on Earth's surface and aiding in climate analysis, ecosystem assessment, and disaster response. In this paper…
Focal Length and Object Pose Estimation via Render and Compare
Georgy Ponimatkin, Yann Labbé, Bryan Russell +2
We introduce FocalPose, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a known…
RANSAC-Flow: generic two-stage image alignment
Xi Shen, François Darmon, Alexei A. Efros +1
This paper considers the generic problem of dense alignment between two images, whether they be two frames of a video, two widely different views of a scene, two paintings depictin…
Online Segmentation of LiDAR Sequences: Dataset and Algorithm
Romain Loiseau, Mathieu Aubry, Loïc Landrieu
Roof-mounted spinning LiDAR sensors are widely used by autonomous vehicles. However, most semantic datasets and algorithms used for LiDAR sequence segmentation operate on $360^\cir…
Learning Joint Surface Atlases
Theo Deprelle, Thibault Groueix, Noam Aigerman +2
This paper describes new techniques for learning atlas-like representations of 3D surfaces, i.e. homeomorphic transformations from a 2D domain to surfaces. Compared to prior work,…
Text region detection in historical astronomical diagrams
Zeynep Sonat Baltacı, Raphaël Baena, Fei Meng +4
Text detection is a crucial task in the analysis of historical documents. While datasets and benchmarks exist for text detection in manuscripts and maps, the study of text in mathe…
Virtual Training for a Real Application: Accurate Object-Robot Relative Localization without Calibration
Vianney Loing, Renaud Marlet, Mathieu Aubry
Localizing an object accurately with respect to a robot is a key step for autonomous robotic manipulation. In this work, we propose to tackle this task knowing only 3D models of th…
Unsupervised Layered Image Decomposition into Object Prototypes
Tom Monnier, Elliot Vincent, Jean Ponce +1
We present an unsupervised learning framework for decomposing images into layers of automatically discovered object models. Contrary to recent approaches that model image layers wi…
Single-view robot pose and joint angle estimation via render & compare
Yann Labbé, Justin Carpentier, Mathieu Aubry +1
We introduce RoboPose, a method to estimate the joint angles and the 6D camera-to-robot pose of a known articulated robot from a single RGB image. This is an important problem to g…
Convolutional Neural Networks for joint object detection and pose estimation: A comparative study
Francisco Massa, Mathieu Aubry, Renaud Marlet
In this paper we study the application of convolutional neural networks for jointly detecting objects depicted in still images and estimating their 3D pose. We identify different f…
Discovering Visual Patterns in Art Collections with Spatially-consistent Feature Learning
Xi Shen, Alexei A. Efros, Mathieu Aubry
Our goal in this paper is to discover near duplicate patterns in large collections of artworks. This is harder than standard instance mining due to differences in the artistic medi…
3D Sketching using Multi-View Deep Volumetric Prediction
Johanna Delanoy, Mathieu Aubry, Phillip Isola +2
Sketch-based modeling strives to bring the ease and immediacy of drawing to the 3D world. However, while drawings are easy for humans to create, they are very challenging for compu…
Deep Exemplar 2D-3D Detection by Adapting from Real to Rendered Views
Francisco Massa, Bryan Russell, Mathieu Aubry
This paper presents an end-to-end convolutional neural network (CNN) for 2D-3D exemplar detection. We demonstrate that the ability to adapt the features of natural images to better…
Monte-Carlo Tree Search for Efficient Visually Guided Rearrangement Planning
Yann Labbé, Sergey Zagoruyko, Igor Kalevatykh +4
We address the problem of visually guided rearrangement planning with many movable objects, i.e., finding a sequence of actions to move a set of objects from an initial arrangement…
Segmenting France Across Four Centuries
Marta López-Rauhut, Hongyu Zhou, Mathieu Aubry +1
Historical maps offer an invaluable perspective into territory evolution across past centuries--long before satellite or remote sensing technologies existed. Deep learning methods…
Improving neural implicit surfaces geometry with patch warping
François Darmon, Bénédicte Bascle, Jean-Clément Devaux +2
Neural implicit surfaces have become an important technique for multi-view 3D reconstruction but their accuracy remains limited. In this paper, we argue that this comes from the di…
Learning Co-segmentation by Segment Swapping for Retrieval and Discovery
Xi Shen, Alexei A. Efros, Armand Joulin +1
The goal of this work is to efficiently identify visually similar patterns in images, e.g. identifying an artwork detail copied between an engraving and an oil painting, or recogni…
Unsupervised cycle-consistent deformation for shape matching
Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2
We propose a self-supervised approach to deep surface deformation. Given a pair of shapes, our algorithm directly predicts a parametric transformation from one shape to the other r…
The Learnable Typewriter: A Generative Approach to Text Analysis
Ioannis Siglidis, Nicolas Gonthier, Julien Gaubil +2
We present a generative document-specific approach to character analysis and recognition in text lines. Our main idea is to build on unsupervised multi-object segmentation methods…
Image Collation: Matching illustrations in manuscripts
Ryad Kaoua, Xi Shen, Alexandra Durr +3
Illustrations are an essential transmission instrument. For an historian, the first step in studying their evolution in a corpus of similar manuscripts is to identify which ones co…
Crafting a multi-task CNN for viewpoint estimation
Francisco Massa, Renaud Marlet, Mathieu Aubry
Convolutional Neural Networks (CNNs) were recently shown to provide state-of-the-art results for object category viewpoint estimation. However different ways of formulating this pr…
Deep Sprite-based Image Models: An Analysis
Zeynep Sonat Baltacı, Romain Loiseau, Mathieu Aubry
While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple problem of identifying recur…
Historical Astronomical Diagrams Decomposition in Geometric Primitives
Syrine Kalleli, Scott Trigg, Ségolène Albouy +2
Automatically extracting the geometric content from the hundreds of thousands of diagrams drawn in historical manuscripts would enable historians to study the diffusion of astronom…
Learning Dense Correspondence via 3D-guided Cycle Consistency
Tinghui Zhou, Philipp Krähenbühl, Mathieu Aubry +2
Discriminative deep learning approaches have shown impressive results for problems where human-labeled ground truth is plentiful, but what about tasks where labels are difficult or…
Diffusion Models as Data Mining Tools
Ioannis Siglidis, Aleksander Holynski, Alexei A. Efros +2
This paper demonstrates how to use generative models trained for image synthesis as tools for visual data mining. Our insight is that since contemporary generative models learn an…
Spherical Perspective on Learning with Normalization Layers
Simon Roburin, Yann de Mont-Marin, Andrei Bursuc +3
Normalization Layers (NLs) are widely used in modern deep-learning architectures. Despite their apparent simplicity, their effect on optimization is not yet fully understood. This…
An Interpretable Deep Learning Approach for Morphological Script Type Analysis
Malamatenia Vlachou-Efstathiou, Ioannis Siglidis, Dominique Stutzmann +1
Defining script types and establishing classification criteria for medieval handwriting is a central aspect of palaeographical analysis. However, existing typologies often encounte…
Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives
Tom Monnier, Jake Austin, Angjoo Kanazawa +2
Given a set of calibrated images of a scene, we present an approach that produces a simple, compact, and actionable 3D world representation by means of 3D primitives. While many ap…
FocalPose++: Focal Length and Object Pose Estimation via Render and Compare
Martin CÃfka, Georgy Ponimatkin, Yann Labbé +4
We introduce FocalPose++, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a kno…
docExtractor: An off-the-shelf historical document element extraction
Tom Monnier, Mathieu Aubry
We present docExtractor, a generic approach for extracting visual elements such as text lines or illustrations from historical documents without requiring any real data annotation.…
Learning elementary structures for 3D shape generation and matching
Theo Deprelle, Thibault Groueix, Matthew Fisher +3
We propose to represent shapes as the deformation and combination of learnable elementary 3D structures, which are primitives resulting from training over a collection of shape. We…
Deep Transformation-Invariant Clustering
Tom Monnier, Thibault Groueix, Mathieu Aubry
Recent advances in image clustering typically focus on learning better deep representations. In contrast, we present an orthogonal approach that does not rely on abstract features…
A Model You Can Hear: Audio Identification with Playable Prototypes
Romain Loiseau, Baptiste Bouvier, Yann Teytaut +3
Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations…
Pixel-wise Agricultural Image Time Series Classification: Comparisons and a Deformable Prototype-based Approach
Elliot Vincent, Jean Ponce, Mathieu Aubry
Improvements in Earth observation by satellites allow for imagery of ever higher temporal and spatial resolution. Leveraging this data for agricultural monitoring is key for addres…
Representing Shape Collections with Alignment-Aware Linear Models
Romain Loiseau, Tom Monnier, Mathieu Aubry +1
In this paper, we revisit the classical representation of 3D point clouds as linear shape models. Our key insight is to leverage deep learning to represent a collection of shapes a…
Share With Thy Neighbors: Single-View Reconstruction by Cross-Instance Consistency
Tom Monnier, Matthew Fisher, Alexei A. Efros +1
Approaches for single-view reconstruction typically rely on viewpoint annotations, silhouettes, the absence of background, multiple views of the same instance, a template shape, or…
Pose from Shape: Deep Pose Estimation for Arbitrary 3D Objects
Yang Xiao, Xuchong Qiu, Pierre-Alain Langlois +2
Most deep pose estimation methods need to be trained for specific object instances or categories. In this work we propose a completely generic deep pose estimation approach, which…
Historical Printed Ornaments: Dataset and Tasks
Sayan Kumar Chaki, Zeynep Sonat Baltaci, Elliot Vincent +5
This paper aims to develop the study of historical printed ornaments with modern unsupervised computer vision. We highlight three complex tasks that are of critical interest to boo…
Large-Scale Historical Watermark Recognition: dataset and a new consistency-based approach
Xi Shen, Ilaria Pastrolin, Oumayma Bounou +4
Historical watermark recognition is a highly practical, yet unsolved challenge for archivists and historians. With a large number of well-defined classes, cluttered and noisy sampl…
Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model
Marta López-Rauhut, Loic Landrieu, Mathieu Aubry +1
New multi-modal large language models (MLLMs) are continuously being trained and deployed, following rapid development cycles. This generative AI frenzy is driving steady increases…
Leveraging Morphology for Historical Script Metrological Analysis
Malamatenia Vlachou Efstathiou, Raphaël Baena, Dominique Stutzmann +1
Advances in handwritten text recognition have enabled large-scale transcription of historical documents, but still provide limited access to interpretable visual measurements for p…
Understanding deep features with computer-generated imagery
Mathieu Aubry, Bryan Russell
We introduce an approach for analyzing the variation of features generated by convolutional neural networks (CNNs) with respect to scene factors that occur in natural images. Such…
Learnable Earth Parser: Discovering 3D Prototypes in Aerial Scans
Romain Loiseau, Elliot Vincent, Mathieu Aubry +1
We propose an unsupervised method for parsing large 3D scans of real-world scenes with easily-interpretable shapes. This work aims to provide a practical tool for analyzing 3D scen…
Learning to Guide Local Feature Matches
François Darmon, Mathieu Aubry, Pascal Monasse
We tackle the problem of finding accurate and robust keypoint correspondences between images. We propose a learning-based approach to guide local feature matches via a learned appr…
CosyPose: Consistent multi-view multi-object 6D pose estimation
Yann Labbé, Justin Carpentier, Mathieu Aubry +1
We introduce an approach for recovering the 6D pose of multiple known objects in a scene captured by a set of input images with unknown camera viewpoints. First, we present a singl…
MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare
Yann Labbé, Lucas Manuelli, Arsalan Mousavian +7
We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a reg…
Impact of base dataset design on few-shot image classification
Othman Sbai, Camille Couprie, Mathieu Aubry
The quality and generality of deep image features is crucially determined by the data they have been trained on, but little is known about this often overlooked effect. In this pap…
Unsupervised Image Decomposition in Vector Layers
Othman Sbai, Camille Couprie, Mathieu Aubry
Deep image generation is becoming a tool to enhance artists and designers creativity potential. In this paper, we aim at making the generation process more structured and easier to…
CoDEx: Combining Domain Expertise for Spatial Generalization in Satellite Image Analysis
Abhishek Kuriyal, Elliot Vincent, Mathieu Aubry +1
Global variations in terrain appearance raise a major challenge for satellite image analysis, leading to poor model performance when training on locations that differ from those en…
Detecting Looted Archaeological Sites from Satellite Image Time Series
Elliot Vincent, Mehraïl Saroufim, Jonathan Chemla +4
Archaeological sites are the physical remains of past human activity and one of the main sources of information about past societies and cultures. However, they are also the target…
Deep Multi-View Stereo gone wild
François Darmon, Bénédicte Bascle, Jean-Clément Devaux +2
Deep multi-view stereo (MVS) methods have been developed and extensively compared on simple datasets, where they now outperform classical approaches. In this paper, we ask whether…