145 citations · 201 across the 10 of their papers we have counts for
6 papers · 1 filter
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Dejia Xu, Weili Nie, Chao Liu +4
Recently video diffusion models have emerged as expressive generative tools for high-quality video content creation readily available to general users. However, these models often…
AGG: Amortized Generative 3D Gaussians for Single Image to 3D
Dejia Xu, Ye Yuan, Morteza Mardani +4
Given the growing need for automatic 3D content creation pipelines, various 3D representations have been studied to generate 3D objects from a single image. Due to its superior ren…
Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models
Jiarui Xu, Sifei Liu, Arash Vahdat +3
We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panopti…
Recurrence without Recurrence: Stable Video Landmark Detection with Deep Equilibrium Models
Paul Micaelli, Arash Vahdat, Hongxu Yin +2
Cascaded computation, whereby predictions are recurrently refined over several stages, has been a persistent theme throughout the development of landmark detection models. In this…
ISB: Image-to-Image Schrödinger Bridge
Guan-Horng Liu, Arash Vahdat, De-An Huang +3
We propose Image-to-Image Schrödinger Bridge (ISB), a new class of conditional diffusion models that directly learn the nonlinear diffusion processes between two given distribu…
Visual Recognition by Counting Instances: A Multi-Instance Cardinality Potential Kernel
Hossein Hajimirsadeghi, Wang Yan, Arash Vahdat +1
Many visual recognition problems can be approached by counting instances. To determine whether an event is present in a long internet video, one could count how many frames seem to…