papers

Publications (41)

cs.CV2022

Pyramid Adversarial Training Improves ViT Performance

Charles Herrmann, Kyle Sargent, Lu Jiang +5

Aggressive data augmentation is a key component of the strong generalization capabilities of Vision Transformer (ViT). One such data augmentation technique is adversarial training…

cs.CV2023

The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation

Saurabh Saxena, Charles Herrmann, Junhwa Hur +4

Denoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity. We show that they also excel in estimating optical flow and…

cs.CV2024

Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov +14

We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthes…

cs.CV2024

DreamWalk: Style Space Exploration using Diffusion Guidance

Michelle Shu, Charles Herrmann, Richard Strong Bowen +2

Text-conditioned diffusion models can generate impressive images, but fall short when it comes to fine-grained control. Unlike direct-editing tools like Photoshop, text conditioned…

cs.CV2024

Boundary Attention: Learning curves, corners, junctions and grouping

Mia Gaia Polansky, Charles Herrmann, Junhwa Hur +3

We present a lightweight network that infers grouping and boundaries, including curves, corners and junctions. It operates in a bottom-up fashion, analogous to classical methods fo…

cs.CV2025

High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion

Junhwa Hur, Charles Herrmann, Saurabh Saxena +6

Despite the recent progress, existing frame interpolation methods still struggle with processing extremely high resolution input and handling challenging cases such as repetitive t…

cs.CV2026

UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images

Junhwa Hur, Charles Herrmann, Songyou Peng +4

Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward model…

cs.CV2023

DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback

Jiao Sun, Deqing Fu, Yushi Hu +8

Despite their wide-spread success, Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user's input text. We introd…

cs.CV2017

A discriminative view of MRF pre-processing algorithms

Chen Wang, Charles Herrmann, Ramin Zabih

While Markov Random Fields (MRFs) are widely used in computer vision, they present a quite challenging inference problem. MRF inference can be accelerated by pre-processing techniq…

cs.CV2022

Kubric: A scalable dataset generator

Klaus Greff, Francois Belletti, Lucas Beyer +32

Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and trainin…

cs.CV2025

Motion Prompting: Controlling Video Generation with Motion Trajectories

Daniel Geng, Charles Herrmann, Junhwa Hur +11

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which stru…

cs.CV2022

Disentangling Architecture and Training for Optical Flow

Deqing Sun, Charles Herrmann, Fitsum Reda +3

How important are training details and datasets to recent optical flow models like RAFT? And do they generalize? To explore these questions, rather than develop a new model, we rev…

cs.CV2026

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

Nate Gillman, Yinghua Zhou, Zitian Tang +6

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying prec…

cs.CV2026

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

Gene Chou, Charles Herrmann, Kyle Genova +6

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produc…

cs.CV2025

MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

Junyi Zhang, Charles Herrmann, Junhwa Hur +5

Estimating geometry from dynamic scenes, where objects move and deform over time, remains a core challenge in computer vision. Current approaches often rely on multi-stage pipeline…

cs.GR2025

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

Zizhang Li, Hong-Xing Yu, Wei Liu +4

WonderPlay is a novel framework integrating physics simulation with video generation for generating action-conditioned dynamic 3D scenes from a single image. While prior works are…

cs.CV2023

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

Saurabh Saxena, Junhwa Hur, Charles Herrmann +2

While methods for monocular depth estimation have made significant strides on standard benchmarks, zero-shot metric depth estimation remains unsolved. Challenges include the joint…

cs.CV2020

Channel selection using Gumbel Softmax

Charles Herrmann, Richard Strong Bowen, Ramin Zabih

Important applications such as mobile computing require reducing the computational costs of neural network inference. Ideally, applications would specify their preferred tradeoff b…

cs.CV2024

WonderJourney: Going from Anywhere to Everywhere

Hong-Xing Yu, Haoyi Duan, Junhwa Hur +8

We introduce WonderJourney, a modularized framework for perpetual 3D scene generation. Unlike prior work on view generation that focuses on a single type of scenes, we start at any…

eess.IV2021

Deep survival analysis with longitudinal X-rays for COVID-19

Michelle Shu, Richard Strong Bowen, Charles Herrmann +3

Time-to-event analysis is an important statistical tool for allocating clinical resources such as ICU beds. However, classical techniques like the Cox model cannot directly incorpo…

cs.CV2024

ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

Kyle Sargent, Zizhang Li, Tanmay Shah +8

We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked b…

cs.CV2023

Accidental Light Probes

Hong-Xing Yu, Samir Agarwala, Charles Herrmann +4

Recovering lighting in a scene from a single image is a fundamental problem in computer vision. While a mirror ball light probe can capture omnidirectional lighting, light probes a…

cs.CV2020

Learning to Autofocus

Charles Herrmann, Richard Strong Bowen, Neal Wadhwa +4

Autofocus is an important task for digital cameras, yet current approaches often exhibit poor performance. We propose a learning-based approach to this problem, and provide a reali…

cs.CV2023

VQ3D: Learning a 3D-Aware Generative Model on ImageNet

Kyle Sargent, Jing Yu Koh, Han Zhang +5

Recent work has shown the possibility of training generative models of 3D content from 2D image collections on small datasets corresponding to a single object class, such as human…

cs.CV2026

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

Leslie Gu, Junhwa Hur, Charles Herrmann +4

GeCo is a geometry-based metric that detects deformation and occlusion inconsistencies in generated videos by combining residual motion and depth cues, providing dense consistency…

#video generation#geometric consistency#motion analysis#depth estimation
cs.CV2020

Object-centered image stitching

Charles Herrmann, Chen Wang, Richard Strong Bowen +2

Image stitching is typically decomposed into three phases: registration, which aligns the source images with a common target image; seam finding, which determines for each target p…

cs.CV2025

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

Kyle Sargent, Ruiqi Gao, Philipp Henzler +5

Evaluations of image compression performance which include human preferences have generally found that naive distortion functions such as MSE are insufficiently aligned to human pe…

cs.CV2025

MASIV: Toward Material-Agnostic System Identification from Videos

Yizhou Zhao, Haoyu Chen, Chunjiang Liu +7

System identification from videos aims to recover object geometry and governing physical laws. Existing methods integrate differentiable rendering with simulation but rely on prede…

cs.CV2021

AutoFlow: Learning a Better Training Set for Optical Flow

Deqing Sun, Daniel Vlasic, Charles Herrmann +6

Synthetic datasets play a critical role in pre-training CNN models for optical flow, but they are painstaking to generate and hard to adapt to new applications. To automate the pro…

cs.CV2025

Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals

Nate Gillman, Charles Herrmann, Michael Freeman +4

Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically m…

cs.CV2025

WonderWorld: Interactive 3D Scene Generation from a Single Image

Hong-Xing Yu, Haoyi Duan, Charles Herrmann +2

We present WonderWorld, a novel framework for interactive 3D scene generation that enables users to interactively specify scene contents and layout and see the created scenes in lo…

cs.LG2023

Substance or Style: What Does Your Image Embedding Know?

Cyrus Rashtchian, Charles Herrmann, Chun-Sung Ferng +5

Probes are small networks that predict properties of underlying data from embeddings, and they provide a targeted, effective way to illuminate the information contained in embeddin…

cs.CV2025

A Simple Approach to Unifying Diffusion-based Conditional Generation

Xirui Li, Charles Herrmann, Kelvin C. K. Chan +4

Recent progress in image generation has sparked research into controlling these models through condition signals, with various methods addressing specific challenges in conditional…

cs.CV2024

Efficient Hybrid Zoom using Camera Fusion on Mobile Phones

Xiaotong Wu, Wei-Sheng Lai, YiChang Shih +4

DSLR cameras can achieve multiple zoom levels via shifting lens distances or swapping lens types. However, these techniques are not possible on smartphone devices due to space cons…

cs.CV2023

A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence

Junyi Zhang, Charles Herrmann, Junhwa Hur +4

Text-to-image diffusion models have made significant advances in generating and editing high-quality images. As a result, numerous approaches have explored the ability of diffusion…

cs.CV2025

MotionV2V: Editing Motion in a Video

Ryan Burgert, Charles Herrmann, Forrester Cole +4

While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explo…

cs.LG2025

Improving realistic semi-supervised learning with doubly robust estimation

Khiem Pham, Charles Herrmann, Ramin Zabih

A major challenge in Semi-Supervised Learning (SSL) is the limited information available about the class distribution in the unlabeled data. In many real-world applications this ar…

cs.CV2020

Robust image stitching with multiple registrations

Charles Herrmann, Chen Wang, Richard Strong Bowen +4

Panorama creation is one of the most widely deployed techniques in computer vision. In addition to industry applications such as Google Street View, it is also used by millions of…

cs.CV2026

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

Junyi Zhang, Charles Herrmann, Junhwa Hur +5

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or lim…

cs.CV2024

Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence

Junyi Zhang, Charles Herrmann, Junhwa Hur +4

While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of insta…

cs.CV2023

Self-supervised AutoFlow

Hsin-Ping Huang, Charles Herrmann, Junhwa Hur +5

Recently, AutoFlow has shown promising results on learning a training set for optical flow, but requires ground truth labels in the target domain to compute its search metric. Obse…