activity
20162024
most citedStyleGAN knows Normal, Depth, Albedo, and More

6 citations · 19 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB

Jae Yong Lee, Yuqun Wu, Chuhang Zou +2

The goal of this paper is to encode a 3D scene into an extremely compact representation from 2D images and to enable its transmittance, decoding and rendering in real-time across v…

cs.CV2024

Anytime Continual Learning for Open Vocabulary Classification

Zhen Zhu, Yiming Gong, Derek Hoiem

We propose an approach for anytime continual learning (AnytimeCL) for open vocabulary image classification. The AnytimeCL problem aims to break away from batch training and rigid m…

cs.CV20233 cited

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Jiasen Lu, Christopher Clark, Sangho Lee +5

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we…

cs.CL20231 cited

WebWISE: Web Interface Control and Sequential Exploration with Large Language Models

Heyi Tao, Sethuraman T, Michal Shlapentokh-Rothman +1

The paper investigates using a Large Language Model (LLM) to automatically perform web software tasks using click, scroll, and text input operations. Previous approaches, such as r…

cs.CV20236 cited

StyleGAN knows Normal, Depth, Albedo, and More

Anand Bhattad, Daniel McKee, Derek Hoiem +1

Intrinsic images, in the original sense, are image-like maps of scene properties like depth, normal, albedo or shading. This paper demonstrates that StyleGAN can easily be induced…

cs.CV20235 cited

Make It So: Steering StyleGAN for Any Image Inversion and Editing

Anand Bhattad, Viraj Shah, Derek Hoiem +1

StyleGAN's disentangled style representation enables powerful image editing by manipulating the latent variables, but accurately mapping real-world images to their latent variables…