papers

Publications (54)

cs.CV2018

3D-CODED : 3D Correspondences by Deep Deformation

Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2

We present a new deep learning approach for matching deformable shapes by introducing {\it Shape Deformation Networks} which jointly encode 3D shapes and correspondences. This is a…

cs.CV2025

General Detection-based Text Line Recognition

Raphael Baena, Syrine Kalleli, Mathieu Aubry

We introduce a general detection-based approach to text line recognition, be it printed (OCR) or handwritten (HTR), with Latin, Chinese, or ciphered characters. Detection-based app…

cs.CV2018

AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation

Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2

We introduce a method for learning to generate the surface of 3D shapes. Our approach represents a 3D shape as a collection of parametric surface elements and, in contrast to metho…

cs.CV2024

Satellite Image Time Series Semantic Change Detection: Novel Architecture and Analysis of Domain Shift

Elliot Vincent, Jean Ponce, Mathieu Aubry

Satellite imagery plays a crucial role in monitoring changes happening on Earth's surface and aiding in climate analysis, ecosystem assessment, and disaster response. In this paper…

cs.CV2022

Focal Length and Object Pose Estimation via Render and Compare

Georgy Ponimatkin, Yann Labbé, Bryan Russell +2

We introduce FocalPose, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a known…

cs.CV2020

RANSAC-Flow: generic two-stage image alignment

Xi Shen, François Darmon, Alexei A. Efros +1

This paper considers the generic problem of dense alignment between two images, whether they be two frames of a video, two widely different views of a scene, two paintings depictin…

cs.CV2022

Online Segmentation of LiDAR Sequences: Dataset and Algorithm

Romain Loiseau, Mathieu Aubry, Loïc Landrieu

Roof-mounted spinning LiDAR sensors are widely used by autonomous vehicles. However, most semantic datasets and algorithms used for LiDAR sequence segmentation operate on $360^\cir…

cs.CG2022

Learning Joint Surface Atlases

Theo Deprelle, Thibault Groueix, Noam Aigerman +2

This paper describes new techniques for learning atlas-like representations of 3D surfaces, i.e. homeomorphic transformations from a 2D domain to surfaces. Compared to prior work,…

cs.CV2026

Text region detection in historical astronomical diagrams

Zeynep Sonat Baltacı, Raphaël Baena, Fei Meng +4

Text detection is a crucial task in the analysis of historical documents. While datasets and benchmarks exist for text detection in manuscripts and maps, the study of text in mathe…

cs.CV2019

Virtual Training for a Real Application: Accurate Object-Robot Relative Localization without Calibration

Vianney Loing, Renaud Marlet, Mathieu Aubry

Localizing an object accurately with respect to a robot is a key step for autonomous robotic manipulation. In this work, we propose to tackle this task knowing only 3D models of th…

cs.CV2021

Unsupervised Layered Image Decomposition into Object Prototypes

Tom Monnier, Elliot Vincent, Jean Ponce +1

We present an unsupervised learning framework for decomposing images into layers of automatically discovered object models. Contrary to recent approaches that model image layers wi…

cs.CV2021

Single-view robot pose and joint angle estimation via render & compare

Yann Labbé, Justin Carpentier, Mathieu Aubry +1

We introduce RoboPose, a method to estimate the joint angles and the 6D camera-to-robot pose of a known articulated robot from a single RGB image. This is an important problem to g…

cs.CV2015

Convolutional Neural Networks for joint object detection and pose estimation: A comparative study

Francisco Massa, Mathieu Aubry, Renaud Marlet

In this paper we study the application of convolutional neural networks for jointly detecting objects depicted in still images and estimating their 3D pose. We identify different f…

cs.CV2019

Discovering Visual Patterns in Art Collections with Spatially-consistent Feature Learning

Xi Shen, Alexei A. Efros, Mathieu Aubry

Our goal in this paper is to discover near duplicate patterns in large collections of artworks. This is harder than standard instance mining due to differences in the artistic medi…

cs.GR2018

3D Sketching using Multi-View Deep Volumetric Prediction

Johanna Delanoy, Mathieu Aubry, Phillip Isola +2

Sketch-based modeling strives to bring the ease and immediacy of drawing to the 3D world. However, while drawings are easy for humans to create, they are very challenging for compu…

cs.CV2016

Deep Exemplar 2D-3D Detection by Adapting from Real to Rendered Views

Francisco Massa, Bryan Russell, Mathieu Aubry

This paper presents an end-to-end convolutional neural network (CNN) for 2D-3D exemplar detection. We demonstrate that the ability to adapt the features of natural images to better…

cs.RO2020

Monte-Carlo Tree Search for Efficient Visually Guided Rearrangement Planning

Yann Labbé, Sergey Zagoruyko, Igor Kalevatykh +4

We address the problem of visually guided rearrangement planning with many movable objects, i.e., finding a sequence of actions to move a set of objects from an initial arrangement…

cs.CV2025

Segmenting France Across Four Centuries

Marta López-Rauhut, Hongyu Zhou, Mathieu Aubry +1

Historical maps offer an invaluable perspective into territory evolution across past centuries--long before satellite or remote sensing technologies existed. Deep learning methods…

cs.CV2022

Improving neural implicit surfaces geometry with patch warping

François Darmon, Bénédicte Bascle, Jean-Clément Devaux +2

Neural implicit surfaces have become an important technique for multi-view 3D reconstruction but their accuracy remains limited. In this paper, we argue that this comes from the di…

cs.CV2022

Learning Co-segmentation by Segment Swapping for Retrieval and Discovery

Xi Shen, Alexei A. Efros, Armand Joulin +1

The goal of this work is to efficiently identify visually similar patterns in images, e.g. identifying an artwork detail copied between an engraving and an oil painting, or recogni…

cs.CV2019

Unsupervised cycle-consistent deformation for shape matching

Thibault Groueix, Matthew Fisher, Vladimir G. Kim +2

We propose a self-supervised approach to deep surface deformation. Given a pair of shapes, our algorithm directly predicts a parametric transformation from one shape to the other r…

cs.CV2023

The Learnable Typewriter: A Generative Approach to Text Analysis

Ioannis Siglidis, Nicolas Gonthier, Julien Gaubil +2

We present a generative document-specific approach to character analysis and recognition in text lines. Our main idea is to build on unsupervised multi-object segmentation methods…

cs.CV2021

Image Collation: Matching illustrations in manuscripts

Ryad Kaoua, Xi Shen, Alexandra Durr +3

Illustrations are an essential transmission instrument. For an historian, the first step in studying their evolution in a corpus of similar manuscripts is to identify which ones co…

cs.CV2016

Crafting a multi-task CNN for viewpoint estimation

Francisco Massa, Renaud Marlet, Mathieu Aubry

Convolutional Neural Networks (CNNs) were recently shown to provide state-of-the-art results for object category viewpoint estimation. However different ways of formulating this pr…

cs.CV2026

Deep Sprite-based Image Models: An Analysis

Zeynep Sonat Baltacı, Romain Loiseau, Mathieu Aubry

While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple problem of identifying recur…

cs.CV2024

Historical Astronomical Diagrams Decomposition in Geometric Primitives

Syrine Kalleli, Scott Trigg, Ségolène Albouy +2

Automatically extracting the geometric content from the hundreds of thousands of diagrams drawn in historical manuscripts would enable historians to study the diffusion of astronom…

cs.CV2016

Learning Dense Correspondence via 3D-guided Cycle Consistency

Tinghui Zhou, Philipp Krähenbühl, Mathieu Aubry +2

Discriminative deep learning approaches have shown impressive results for problems where human-labeled ground truth is plentiful, but what about tasks where labels are difficult or…

cs.CV2024

Diffusion Models as Data Mining Tools

Ioannis Siglidis, Aleksander Holynski, Alexei A. Efros +2

This paper demonstrates how to use generative models trained for image synthesis as tools for visual data mining. Our insight is that since contemporary generative models learn an…

cs.LG2022

Spherical Perspective on Learning with Normalization Layers

Simon Roburin, Yann de Mont-Marin, Andrei Bursuc +3

Normalization Layers (NLs) are widely used in modern deep-learning architectures. Despite their apparent simplicity, their effect on optimization is not yet fully understood. This…

cs.CV2024

An Interpretable Deep Learning Approach for Morphological Script Type Analysis

Malamatenia Vlachou-Efstathiou, Ioannis Siglidis, Dominique Stutzmann +1

Defining script types and establishing classification criteria for medieval handwriting is a central aspect of palaeographical analysis. However, existing typologies often encounte…

cs.CV2023

Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives

Tom Monnier, Jake Austin, Angjoo Kanazawa +2

Given a set of calibrated images of a scene, we present an approach that produces a simple, compact, and actionable 3D world representation by means of 3D primitives. While many ap…

cs.CV2024

FocalPose++: Focal Length and Object Pose Estimation via Render and Compare

Martin Cífka, Georgy Ponimatkin, Yann Labbé +4

We introduce FocalPose++, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a kno…

cs.CV2020

docExtractor: An off-the-shelf historical document element extraction

Tom Monnier, Mathieu Aubry

We present docExtractor, a generic approach for extracting visual elements such as text lines or illustrations from historical documents without requiring any real data annotation.…

cs.CV2019

Learning elementary structures for 3D shape generation and matching

Theo Deprelle, Thibault Groueix, Matthew Fisher +3

We propose to represent shapes as the deformation and combination of learnable elementary 3D structures, which are primitives resulting from training over a collection of shape. We…

cs.CV2020

Deep Transformation-Invariant Clustering

Tom Monnier, Thibault Groueix, Mathieu Aubry

Recent advances in image clustering typically focus on learning better deep representations. In contrast, we present an orthogonal approach that does not rely on abstract features…

cs.SD2022

A Model You Can Hear: Audio Identification with Playable Prototypes

Romain Loiseau, Baptiste Bouvier, Yann Teytaut +3

Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations…

cs.CV2024

Pixel-wise Agricultural Image Time Series Classification: Comparisons and a Deformable Prototype-based Approach

Elliot Vincent, Jean Ponce, Mathieu Aubry

Improvements in Earth observation by satellites allow for imagery of ever higher temporal and spatial resolution. Leveraging this data for agricultural monitoring is key for addres…

cs.CV2021

Representing Shape Collections with Alignment-Aware Linear Models

Romain Loiseau, Tom Monnier, Mathieu Aubry +1

In this paper, we revisit the classical representation of 3D point clouds as linear shape models. Our key insight is to leverage deep learning to represent a collection of shapes a…

cs.CV2022

Share With Thy Neighbors: Single-View Reconstruction by Cross-Instance Consistency

Tom Monnier, Matthew Fisher, Alexei A. Efros +1

Approaches for single-view reconstruction typically rely on viewpoint annotations, silhouettes, the absence of background, multiple views of the same instance, a template shape, or…

cs.CV2019

Pose from Shape: Deep Pose Estimation for Arbitrary 3D Objects

Yang Xiao, Xuchong Qiu, Pierre-Alain Langlois +2

Most deep pose estimation methods need to be trained for specific object instances or categories. In this work we propose a completely generic deep pose estimation approach, which…

cs.CV2024

Historical Printed Ornaments: Dataset and Tasks

Sayan Kumar Chaki, Zeynep Sonat Baltaci, Elliot Vincent +5

This paper aims to develop the study of historical printed ornaments with modern unsupervised computer vision. We highlight three complex tasks that are of critical interest to boo…

cs.CV2019

Large-Scale Historical Watermark Recognition: dataset and a new consistency-based approach

Xi Shen, Ilaria Pastrolin, Oumayma Bounou +4

Historical watermark recognition is a highly practical, yet unsolved challenge for archivists and historians. With a large number of well-defined classes, cluttered and noisy sampl…

cs.AI2026

Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model

Marta López-Rauhut, Loic Landrieu, Mathieu Aubry +1

New multi-modal large language models (MLLMs) are continuously being trained and deployed, following rapid development cycles. This generative AI frenzy is driving steady increases…

cs.CV2026

Leveraging Morphology for Historical Script Metrological Analysis

Malamatenia Vlachou Efstathiou, Raphaël Baena, Dominique Stutzmann +1

Advances in handwritten text recognition have enabled large-scale transcription of historical documents, but still provide limited access to interpretable visual measurements for p…

cs.CV2015

Understanding deep features with computer-generated imagery

Mathieu Aubry, Bryan Russell

We introduce an approach for analyzing the variation of features generated by convolutional neural networks (CNNs) with respect to scene factors that occur in natural images. Such…

cs.CV2024

Learnable Earth Parser: Discovering 3D Prototypes in Aerial Scans

Romain Loiseau, Elliot Vincent, Mathieu Aubry +1

We propose an unsupervised method for parsing large 3D scans of real-world scenes with easily-interpretable shapes. This work aims to provide a practical tool for analyzing 3D scen…

cs.CV2020

Learning to Guide Local Feature Matches

François Darmon, Mathieu Aubry, Pascal Monasse

We tackle the problem of finding accurate and robust keypoint correspondences between images. We propose a learning-based approach to guide local feature matches via a learned appr…

cs.CV2020

CosyPose: Consistent multi-view multi-object 6D pose estimation

Yann Labbé, Justin Carpentier, Mathieu Aubry +1

We introduce an approach for recovering the 6D pose of multiple known objects in a scene captured by a set of input images with unknown camera viewpoints. First, we present a singl…

cs.CV2022

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

Yann Labbé, Lucas Manuelli, Arsalan Mousavian +7

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a reg…

cs.CV2020

Impact of base dataset design on few-shot image classification

Othman Sbai, Camille Couprie, Mathieu Aubry

The quality and generality of deep image features is crucially determined by the data they have been trained on, but little is known about this often overlooked effect. In this pap…

cs.CV2019

Unsupervised Image Decomposition in Vector Layers

Othman Sbai, Camille Couprie, Mathieu Aubry

Deep image generation is becoming a tool to enhance artists and designers creativity potential. In this paper, we aim at making the generation process more structured and easier to…

cs.CV2025

CoDEx: Combining Domain Expertise for Spatial Generalization in Satellite Image Analysis

Abhishek Kuriyal, Elliot Vincent, Mathieu Aubry +1

Global variations in terrain appearance raise a major challenge for satellite image analysis, leading to poor model performance when training on locations that differ from those en…

cs.CV2024

Detecting Looted Archaeological Sites from Satellite Image Time Series

Elliot Vincent, Mehraïl Saroufim, Jonathan Chemla +4

Archaeological sites are the physical remains of past human activity and one of the main sources of information about past societies and cultures. However, they are also the target…

cs.CV2021

Deep Multi-View Stereo gone wild

François Darmon, Bénédicte Bascle, Jean-Clément Devaux +2

Deep multi-view stereo (MVS) methods have been developed and extensively compared on simple datasets, where they now outperform classical approaches. In this paper, we ask whether…