papers

Publications (15)

cs.CV2025

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

Guo Chen, Zhiqi Li, Shihao Wang +16

We introduce Eagle 2.5, a family of frontier vision-language models (VLMs) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and h…

cs.CV2025

SIEDD: Shared-Implicit Encoder with Discrete Decoders

Vikram Rangarajan, Shishira Maiya, Max Ehrlich +1

Implicit Neural Representations (INRs) offer exceptional fidelity for video compression by learning per-video optimized functions, but their adoption is crippled by impractically s…

cs.CV2024

Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics

Shishira R Maiya, Anubhav Gupta, Matthew Gwilliam +2

Implicit Neural Networks (INRs) have emerged as powerful representations to encode all forms of data, including images, videos, audios, and scenes. With video, many INRs for video…

cs.CV2021

A Frequency Perspective of Adversarial Robustness

Shishira R Maiya, Max Ehrlich, Vatsal Agarwal +3

Adversarial examples pose a unique challenge for deep learning systems. Despite recent advances in both attacks and defenses, there is still a lack of clarity and consensus in the…

cs.LG2025

Wolf: Dense Video Captioning with a World Summarization Framework

Boyi Li, Ligeng Zhu, Ran Tian +20

We propose Wolf, a WOrLd summarization Framework for accurate video captioning. Wolf is an automated captioning framework that adopts a mixture-of-experts approach, leveraging comp…

eess.IV2022

ReLaX: Retinal Layer Attribution for Guided Explanations of Automated Optical Coherence Tomography Classification

Evan Wen, Rebecca Sorenson, Max Ehrlich

30 million Optical Coherence Tomography (OCT) imaging tests are issued annually to diagnose various retinal diseases, but accurate diagnosis of OCT scans requires trained eye care…

cs.CV2021

Unsupervised Super-Resolution of Satellite Imagery for High Fidelity Material Label Transfer

Arthita Ghosh, Max Ehrlich, Larry Davis +1

Urban material recognition in remote sensing imagery is a highly relevant, yet extremely challenging problem due to the difficulty of obtaining human annotations, especially on low…

eess.IV2022

The First Principles of Deep Learning and Compression

Max Ehrlich

The deep learning revolution incited by the 2012 Alexnet paper has been transformative for the field of computer vision. Many problems which were severely limited using classical s…

cs.LG2019

Deep Residual Learning in the JPEG Transform Domain

Max Ehrlich, Larry Davis

We introduce a general method of performing Residual Network inference and learning in the JPEG transform domain that allows the network to consume compressed images as input. Our…

cs.CV2022

NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-wise Modeling

Shishira R Maiya, Sharath Girish, Max Ehrlich +6

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exp…

cs.CV2021

Analyzing and Mitigating JPEG Compression Defects in Deep Learning

Max Ehrlich, Larry Davis, Ser-Nam Lim +1

With the proliferation of deep learning methods, many computer vision problems which were considered academic are now viable in the consumer setting. One drawback of consumer appli…

eess.IV2020

Quantization Guided JPEG Artifact Correction

Max Ehrlich, Larry Davis, Ser-Nam Lim +1

The JPEG image compression algorithm is the most popular method of image compression because of its ability for large compression ratios. However, to achieve such high compression,…

cs.CV2024

Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions

Namitha Padmanabhan, Matthew Gwilliam, Pulkit Kumar +3

The many variations of Implicit Neural Representations (INRs), where a neural network is trained as a continuous representation of a signal, have tremendous practical utility for d…

eess.IV2023

Leveraging Bitstream Metadata for Fast, Accurate, Generalized Compressed Video Quality Enhancement

Max Ehrlich, Jon Barker, Namitha Padmanabhan +4

Video compression is a central feature of the modern internet powering technologies from social media to video conferencing. While video compression continues to mature, for many c…

cs.CV2016

Action-Affect Classification and Morphing using Multi-Task Representation Learning

Timothy J. Shields, Mohamed R. Amer, Max Ehrlich +1

Most recent work focused on affect from facial expressions, and not as much on body. This work focuses on body affect analysis. Affect does not occur in isolation. Humans usually c…