Publications (29)
Confusing Deep Convolution Networks by Relabelling
Leigh Robinson, Benjamin Graham
Deep convolutional neural networks have become the gold standard for image recognition tasks, demonstrating many current state-of-the-art results and even achieving near-human leve…
Exploring Data-Efficient 3D Scene Understanding with Contrastive Scene Contexts
Ji Hou, Benjamin Graham, Matthias NieÃner +1
The rapid progress in 3D scene understanding has come with growing demand for data; however, collecting and annotating 3D scenes (e.g. point clouds) are notoriously hard. For examp…
DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third Dimension
Roman Shapovalov, David Novotny, Benjamin Graham +2
We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. We contribute DensePose 3D, a method that can learn such reconstructions in a w…
Real-time volumetric rendering of dynamic humans
Ignacio Rocco, Iurii Makarov, Filippos Kokkinos +4
We present a method for fast 3D reconstruction and real-time rendering of dynamic humans from monocular videos with accompanying parametric body fits. Our method can reconstruct a…
Pri3D: Can 3D Priors Help 2D Representation Learning?
Ji Hou, Saining Xie, Benjamin Graham +2
Recent advances in 3D perception have shown impressive progress in understanding geometric structures of 3Dshapes and even scenes. Inspired by these advances in geometric understan…
RidgeSfM: Structure from Motion via Robust Pairwise Matching Under Depth Uncertainty
Benjamin Graham, David Novotny
We consider the problem of simultaneously estimating a dense depth map and camera pose for a large set of images of an indoor scene. While classical SfM pipelines rely on a two-ste…
C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From Motion
David Novotny, Nikhila Ravi, Benjamin Graham +2
We propose C3DPO, a method for extracting 3D models of deformable objects from 2D keypoint annotations in unconstrained images. We do so by learning a deep network that reconstruct…
The iisignature library: efficient calculation of iterated-integral signatures and log signatures
Jeremy Reizenstein, Benjamin Graham
Iterated-integral signatures and log signatures are vectors calculated from a path that characterise its shape. They come from the theory of differential equations driven by rough…
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham +3
Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transitions that dominate downstream p…
Self-Supervised Correspondence Estimation via Multiview Registration
Mohamed El Banani, Ignacio Rocco, David Novotny +4
Video provides us with the spatio-temporal consistency needed for visual learning. Recent approaches have utilized this signal to learn correspondence estimation from close-by fram…
Unsupervised learning with sparse space-and-time autoencoders
Benjamin Graham
We use spatially-sparse two, three and four dimensional convolutional autoencoder networks to model sparse structures in 2D space, 3D space, and 3+1=4 dimensional space-time. We ev…
Unsupervised 2D-3D lifting of non-rigid objects using local constraints
Shalini Maiti, Lourdes Agapito, Benjamin Graham
For non-rigid objects, predicting the 3D shape from 2D keypoint observations is ill-posed due to occlusions, and the need to disentangle changes in viewpoint and changes in shape.…
Equi-normalization of Neural Networks
Pierre Stock, Benjamin Graham, Rémi Gribonval +1
Modern neural networks are over-parametrized. In particular, each rectified linear hidden unit can be modified by a multiplicative factor by adjusting input and output weights, wit…
A binary deletion channel with a fixed number of deletions
Benjamin Graham
Suppose a binary string x = x_1...x_n is being broadcast repeatedly over a faulty communication channel. Each time, the channel delivers a fixed number m of the digits (m<n) with t…
Large-Scale 3D Shape Reconstruction and Segmentation from ShapeNet Core55
Li Yi, Lin Shao, Manolis Savva +47
We introduce a large-scale 3D shape understanding benchmark using data and annotation from ShapeNet 3D object database. The benchmark consists of two tasks: part-level segmentation…
3D Semantic Segmentation with Submanifold Sparse Convolutional Networks
Benjamin Graham, Martin Engelcke, Laurens van der Maaten
Convolutional networks are the de-facto standard for analyzing spatio-temporal data such as images, videos, and 3D shapes. Whilst some of this data is naturally dense (e.g., photos…
Meta 3D Gen
Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17
We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…
Fractional Max-Pooling
Benjamin Graham
Convolutional networks almost always incorporate some form of spatial pooling, and very often it is alpha times alpha max-pooling with alpha=2. Max-pooling act on the hidden layers…
Spatially-sparse convolutional neural networks
Benjamin Graham
Convolutional neural networks (CNNs) perform well on problems such as handwriting recognition and image classification. However, the performance of the networks is often limited by…
CoTracker: It is Better to Track Together
Nikita Karaev, Ignacio Rocco, Benjamin Graham +3
We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points indep…
DynamicStereo: Consistent Dynamic Depth from Stereo Videos
Nikita Karaev, Ignacio Rocco, Benjamin Graham +3
We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, le…
And the Bit Goes Down: Revisiting the Quantization of Neural Networks
Pierre Stock, Armand Joulin, Rémi Gribonval +2
In this paper, we address the problem of reducing the memory footprint of convolutional network architectures. We introduce a vector quantization method that aims at preserving the…
Sharp thresholds for the random-cluster and Ising models
Benjamin Graham, Geoffrey Grimmett
A sharp-threshold theorem is proved for box-crossing probabilities on the square lattice. The models in question are the random-cluster model near the self-dual point $p_{\mathrm {…
Training with Quantization Noise for Extreme Model Compression
Angela Fan, Pierre Stock, Benjamin Graham +4
We tackle the problem of producing compact models, maximizing their accuracy for a given model size. A standard solution is to train networks with Quantization Aware Training, wher…
Low-Precision Batch-Normalized Activations
Benjamin Graham
Artificial neural networks can be trained with relatively low-precision floating-point and fixed-point arithmetic, using between one and 16 bits. Previous works have focused on rel…
Submanifold Sparse Convolutional Networks
Benjamin Graham, Laurens van der Maaten
Convolutional network are the de-facto standard for analysing spatio-temporal data such as images, videos, 3D shapes, etc. Whilst some of this data is naturally dense (for instance…
3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data
Benjamin Biggs, Sébastien Ehrhadt, Hanbyul Joo +3
We consider the problem of obtaining dense 3D reconstructions of humans from single and partially occluded views. In such cases, the visual evidence is usually insufficient to iden…
Sparse arrays of signatures for online character recognition
Benjamin Graham
In mathematics the signature of a path is a collection of iterated integrals, commonly used for solving differential equations. We show that the path signature, used as a set of fe…
A Deep Learning Based Method for Fast Registration of Cardiac Magnetic Resonance Images
Benjamin Graham
Image registration is used in many medical image analysis applications, such as tracking the motion of tissue in cardiac images, where cardiac kinematics can be an indicator of tis…