activity
20222024
most citedZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

175 citations · 287 across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV20241 cited

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

Abdelrahman Eldesokey, Peter Wonka

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of…

cs.CV20241 cited

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

Bing Li, Cheng Zheng, Wenxuan Zhu +4

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new c…

cs.CV2024

PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation

Zhenyu Li, Shariq Farooq Bhat, Peter Wonka

This paper introduces PatchRefiner, an advanced framework for metric single image depth estimation aimed at high-resolution real-domain inputs. While depth estimation is crucial fo…

cs.CV2024

LASPA: Latent Spatial Alignment for Fast Training-free Single Image Editing

Yazeed Alharbi, Peter Wonka

We present a novel, training-free approach for textual editing of real images using diffusion models. Unlike prior methods that rely on computationally expensive finetuning, our ap…

cs.CV2024

AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning

Wamiq Reyaz Para, Abdelrahman Eldesokey, Zhenyu Li +3

We introduce an approach for 3D head avatar generation and editing with multi-modal conditioning based on a 3D Generative Adversarial Network (GAN) and a Latent Diffusion Model (LD…

cs.CV2024

Deep Learning-based Image and Video Inpainting: A Survey

Weize Quan, Jiaxi Chen, Yanli Liu +2

Image and video inpainting is a classic problem in computer vision and computer graphics, aiming to fill in the plausible and realistic content in the missing areas of images and v…