papers

Publications (29)

cs.CV2015

Confusing Deep Convolution Networks by Relabelling

Leigh Robinson, Benjamin Graham

Deep convolutional neural networks have become the gold standard for image recognition tasks, demonstrating many current state-of-the-art results and even achieving near-human leve…

cs.CV2021

Exploring Data-Efficient 3D Scene Understanding with Contrastive Scene Contexts

Ji Hou, Benjamin Graham, Matthias Nießner +1

The rapid progress in 3D scene understanding has come with growing demand for data; however, collecting and annotating 3D scenes (e.g. point clouds) are notoriously hard. For examp…

cs.CV2021

DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third Dimension

Roman Shapovalov, David Novotny, Benjamin Graham +2

We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. We contribute DensePose 3D, a method that can learn such reconstructions in a w…

cs.CV2023

Real-time volumetric rendering of dynamic humans

Ignacio Rocco, Iurii Makarov, Filippos Kokkinos +4

We present a method for fast 3D reconstruction and real-time rendering of dynamic humans from monocular videos with accompanying parametric body fits. Our method can reconstruct a…

cs.CV2021

Pri3D: Can 3D Priors Help 2D Representation Learning?

Ji Hou, Saining Xie, Benjamin Graham +2

Recent advances in 3D perception have shown impressive progress in understanding geometric structures of 3Dshapes and even scenes. Inspired by these advances in geometric understan…

cs.CV2020

RidgeSfM: Structure from Motion via Robust Pairwise Matching Under Depth Uncertainty

Benjamin Graham, David Novotny

We consider the problem of simultaneously estimating a dense depth map and camera pose for a large set of images of an indoor scene. While classical SfM pipelines rely on a two-ste…

cs.CV2019

C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From Motion

David Novotny, Nikhila Ravi, Benjamin Graham +2

We propose C3DPO, a method for extracting 3D models of deformable objects from 2D keypoint annotations in unconstrained images. We do so by learning a deep network that reconstruct…

cs.DS2018

The iisignature library: efficient calculation of iterated-integral signatures and log signatures

Jeremy Reizenstein, Benjamin Graham

Iterated-integral signatures and log signatures are vectors calculated from a path that characterise its shape. They come from the theory of differential equations driven by rough…

cs.LG2026

PROWL: Prioritized Regret-Driven Optimization for World Model Learning

Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham +3

Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transitions that dominate downstream p…

cs.CV2022

Self-Supervised Correspondence Estimation via Multiview Registration

Mohamed El Banani, Ignacio Rocco, David Novotny +4

Video provides us with the spatio-temporal consistency needed for visual learning. Recent approaches have utilized this signal to learn correspondence estimation from close-by fram…

cs.CV2018

Unsupervised learning with sparse space-and-time autoencoders

Benjamin Graham

We use spatially-sparse two, three and four dimensional convolutional autoencoder networks to model sparse structures in 2D space, 3D space, and 3+1=4 dimensional space-time. We ev…

cs.CV2025

Unsupervised 2D-3D lifting of non-rigid objects using local constraints

Shalini Maiti, Lourdes Agapito, Benjamin Graham

For non-rigid objects, predicting the 3D shape from 2D keypoint observations is ill-posed due to occlusions, and the need to disentangle changes in viewpoint and changes in shape.…

cs.CV2019

Equi-normalization of Neural Networks

Pierre Stock, Benjamin Graham, Rémi Gribonval +1

Modern neural networks are over-parametrized. In particular, each rectified linear hidden unit can be modified by a multiplicative factor by adjusting input and output weights, wit…

math.PR2013

A binary deletion channel with a fixed number of deletions

Benjamin Graham

Suppose a binary string x = x_1...x_n is being broadcast repeatedly over a faulty communication channel. Each time, the channel delivers a fixed number m of the digits (m<n) with t…

cs.CV2017

Large-Scale 3D Shape Reconstruction and Segmentation from ShapeNet Core55

Li Yi, Lin Shao, Manolis Savva +47

We introduce a large-scale 3D shape understanding benchmark using data and annotation from ShapeNet 3D object database. The benchmark consists of two tasks: part-level segmentation…

cs.CV2017

3D Semantic Segmentation with Submanifold Sparse Convolutional Networks

Benjamin Graham, Martin Engelcke, Laurens van der Maaten

Convolutional networks are the de-facto standard for analyzing spatio-temporal data such as images, videos, and 3D shapes. Whilst some of this data is naturally dense (e.g., photos…

cs.CV2024

Meta 3D Gen

Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17

We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…

cs.CV2015

Fractional Max-Pooling

Benjamin Graham

Convolutional networks almost always incorporate some form of spatial pooling, and very often it is alpha times alpha max-pooling with alpha=2. Max-pooling act on the hidden layers…

cs.CV2014

Spatially-sparse convolutional neural networks

Benjamin Graham

Convolutional neural networks (CNNs) perform well on problems such as handwriting recognition and image classification. However, the performance of the networks is often limited by…

cs.CV2024

CoTracker: It is Better to Track Together

Nikita Karaev, Ignacio Rocco, Benjamin Graham +3

We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points indep…

cs.CV2023

DynamicStereo: Consistent Dynamic Depth from Stereo Videos

Nikita Karaev, Ignacio Rocco, Benjamin Graham +3

We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, le…

cs.CV2020

And the Bit Goes Down: Revisiting the Quantization of Neural Networks

Pierre Stock, Armand Joulin, Rémi Gribonval +2

In this paper, we address the problem of reducing the memory footprint of convolutional network architectures. We introduce a vector quantization method that aims at preserving the…

math.PR2011

Sharp thresholds for the random-cluster and Ising models

Benjamin Graham, Geoffrey Grimmett

A sharp-threshold theorem is proved for box-crossing probabilities on the square lattice. The models in question are the random-cluster model near the self-dual point $p_{\mathrm {…

cs.LG2021

Training with Quantization Noise for Extreme Model Compression

Angela Fan, Pierre Stock, Benjamin Graham +4

We tackle the problem of producing compact models, maximizing their accuracy for a given model size. A standard solution is to train networks with Quantization Aware Training, wher…

cs.NE2017

Low-Precision Batch-Normalized Activations

Benjamin Graham

Artificial neural networks can be trained with relatively low-precision floating-point and fixed-point arithmetic, using between one and 16 bits. Previous works have focused on rel…

cs.NE2017

Submanifold Sparse Convolutional Networks

Benjamin Graham, Laurens van der Maaten

Convolutional network are the de-facto standard for analysing spatio-temporal data such as images, videos, 3D shapes, etc. Whilst some of this data is naturally dense (for instance…

cs.CV2020

3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data

Benjamin Biggs, Sébastien Ehrhadt, Hanbyul Joo +3

We consider the problem of obtaining dense 3D reconstructions of humans from single and partially occluded views. In such cases, the visual evidence is usually insufficient to iden…

cs.CV2013

Sparse arrays of signatures for online character recognition

Benjamin Graham

In mathematics the signature of a path is a collection of iterated integrals, commonly used for solving differential equations. We show that the path signature, used as a set of fe…

eess.IV2025

A Deep Learning Based Method for Fast Registration of Cardiac Magnetic Resonance Images

Benjamin Graham

Image registration is used in many medical image analysis applications, such as tracking the motion of tissue in cardiac images, where cardiac kinematics can be an indicator of tis…