papers

Publications (22)

cs.CV2020

End-to-end Learning of Compressible Features

Saurabh Singh, Sami Abu-El-Haija, Nick Johnston +3

Pre-trained convolutional neural networks (CNNs) are powerful off-the-shelf feature generators and have been shown to perform very well on a variety of tasks. Unfortunately, the ge…

cs.GR2021

LVAC: Learned Volumetric Attribute Compression for Point Clouds using Coordinate Based Networks

Berivan Isik, Philip A. Chou, Sung Jin Hwang +2

We consider the attributes of a point cloud as samples of a vector-valued volumetric function at discrete positions. To compress the attributes given the positions, we compress the…

eess.IV2020

High-Fidelity Generative Image Compression

Fabian Mentzer, George Toderici, Michael Tschannen +1

We extensively study how to combine Generative Adversarial Networks and learned compression to obtain a state-of-the-art generative lossy compression system. In particular, we inve…

cs.CV2022

VCT: A Video Compression Transformer

Fabian Mentzer, George Toderici, David Minnen +4

We show how transformers can be used to vastly simplify neural video compression. Previous methods have been relying on an increasing number of architectural biases and priors, inc…

cs.CV2018

AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions

Chunhui Gu, Chen Sun, David A. Ross +9

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video…

cs.CV2015

Efficient Large Scale Video Classification

Balakrishnan Varadarajan, George Toderici, Sudheendra Vijayanarasimhan +1

Video classification has advanced tremendously over the recent years. A large part of the improvements in video classification had to do with the work done by the image classificat…

cs.CV2018

Joint Autoregressive and Hierarchical Priors for Learned Image Compression

David Minnen, Johannes Ballé, George Toderici

Recent models for learned image compression are based on autoencoders, learning approximately invertible mappings from pixels to a quantized latent representation. These are combin…

cs.CV2017

Improved Lossy Image Compression with Priming and Spatially Adaptive Bit Rates for Recurrent Networks

Nick Johnston, Damien Vincent, David Minnen +6

We propose a method for lossy image compression based on recurrent, convolutional neural networks that outperforms BPG (4:2:0 ), WebP, JPEG2000, and JPEG as measured by MS-SSIM. We…

cs.CV2016

YouTube-8M: A Large-Scale Video Classification Benchmark

Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee +4

Many recent advancements in Computer Vision are attributed to large datasets. Open-source software packages for Machine Learning and inexpensive commodity hardware have reduced the…

cs.CV2023

Multi-Realism Image Compression with a Conditional Generator

Eirikur Agustsson, David Minnen, George Toderici +1

By optimizing the rate-distortion-realism trade-off, generative compression approaches produce detailed, realistic images, even at low bit rates, instead of the blurry reconstructi…

cs.CV2016

Variable Rate Image Compression with Recurrent Neural Networks

George Toderici, Sean M. O'Malley, Sung Jin Hwang +5

A large fraction of Internet traffic is now driven by requests from mobile devices with relatively small screens and often stringent bandwidth requirements. Due to these factors, i…

cs.CV2018

Towards a Semantic Perceptual Image Metric

Troy Chinen, Johannes Ballé, Chunhui Gu +8

We present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 14…

cs.CV2018

Image-Dependent Local Entropy Models for Learned Image Compression

David Minnen, George Toderici, Saurabh Singh +2

The leading approach for image compression with artificial neural networks (ANNs) is to learn a nonlinear transform and a fixed entropy model that are optimized for rate-distortion…

cs.CV2025

Towards flexible perception with visual memory

Robert Geirhos, Priyank Jaini, Austin Stone +5

Training a neural network is a monolithic endeavor, akin to carving knowledge into stone: once the process is completed, editing the knowledge in a network is hard, since all infor…

cs.IT2020

Nonlinear Transform Coding

Johannes Ballé, Philip A. Chou, David Minnen +5

We review a class of methods that can be collected under the name nonlinear transform coding (NTC), which over the past few years have become competitive with the best linear trans…

cs.CV2015

Beyond Short Snippets: Deep Networks for Video Classification

Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan +3

Convolutional neural networks (CNNs) have been extensively applied for image recognition problems giving state-of-the-art results on recognition, detection, segmentation and retrie…

eess.IV2024

High-Fidelity Image Compression with Score-based Generative Models

Emiel Hoogeboom, Eirikur Agustsson, Fabian Mentzer +3

Despite the tremendous success of diffusion generative models in text-to-image generation, replicating this success in the domain of image compression has proven difficult. In this…

cs.CV2017

Full Resolution Image Compression with Recurrent Neural Networks

George Toderici, Damien Vincent, Nick Johnston +4

This paper presents a set of full-resolution lossy image compression methods based on neural networks. Each of the architectures we describe can provide variable compression rates…

eess.IV2022

Neural Video Compression using GANs for Detail Synthesis and Propagation

Fabian Mentzer, Eirikur Agustsson, Johannes Ballé +3

We present the first neural video compression method based on generative adversarial networks (GANs). Our approach significantly outperforms previous neural and non-neural video co…

cs.CV2017

Target-Quality Image Compression with Recurrent, Convolutional Neural Networks

Michele Covell, Nick Johnston, David Minnen +5

We introduce a stop-code tolerant (SCT) approach to training recurrent convolutional neural networks for lossy image compression. Our methods introduce a multi-pass training method…

cs.CV2018

Spatially adaptive image compression using a tiled deep network

David Minnen, George Toderici, Michele Covell +6

Deep neural networks represent a powerful class of function approximators that can learn to compress and reconstruct images. Existing image compression algorithms based on neural n…

cs.CV2015

Pose Embeddings: A Deep Architecture for Learning to Match Human Poses

Greg Mori, Caroline Pantofaru, Nisarg Kothari +4

We present a method for learning an embedding that places images of humans in similar poses nearby. This embedding can be used as a direct method of comparing images based on human…