From the 1 of 18 linked papers with an AI index.
5 papers · 2 filters
Repurposing 2D Diffusion Models for 3D Shape Completion
Yao He, Youngjoong Kwon, Tiange Xiang +2
We present a framework that adapts 2D diffusion models for 3D shape completion from incomplete point clouds. While text-to-image diffusion models have achieved remarkable success w…
T*: Re-thinking Temporal Search for Long-Form Video Understanding
Jinhui Ye, Zihan Wang, Haosen Sun +9
Efficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding…
Wild2Avatar: Rendering Humans Behind Occlusions
Tiange Xiang, Adam Sun, Scott Delp +3
Rendering the visual appearance of moving humans from occluded monocular videos is a challenging task. Most existing research renders 3D humans under ideal conditions, requiring a…
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
Tiange Xiang, Kai Li, Chengjiang Long +5
Recent advances in text-to-image diffusion models have been driven by the increasing availability of paired 2D data. However, the development of 3D diffusion models has been hinder…
Towards Fine-Grained Video Question Answering
Wei Dai, Alan Luo, Zane Durante +5
In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial gra…