7 papers
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
Jeong-gi Kwak, Sho Kagami, Yuki Ono +1
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry est…
SONIC: Spectral Optimization of Noise for Inpainting with Consistency
Seungyeon Baek, Erqun Dong, Shadan Namazifard +2
We propose a novel training-free method for inpainting with off-the-shelf text-to-image models. While guidance-based methods in theory allow generic models to be used for inverse p…
FullCircle: Effortless 3D Reconstruction from Casual 360 Captures
Yalda Foroutan, Ipek Oztas, Daniel Rebain +4
Radiance fields have emerged as powerful tools for 3D scene reconstruction. However, casual capture remains challenging due to the narrow field of view of perspective cameras, whic…
Making Video Models Adhere to User Intent with Minor Adjustments
Daniel Ajisafe, Eric Hedlin, Helge Rhodin +1
With the recent drastic advancements in text-to-video diffusion models, controlling their generations has drawn interest. A popular way for control is through bounding boxes or lay…
ROODI: Reconstructing Occluded Objects with Denoising Inpainters
Yeonjin Chang, Erqun Dong, Seunghyeon Seo +2
While the quality of novel-view images has improved dramatically with 3D Gaussian Splatting, extracting specific objects from scenes remains challenging. Isolating individual 3D Ga…
HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories
Eric Hedlin, Munawar Hayat, Fatih Porikli +2
To efficiently adapt large models or to train generative models of neural representations, Hypernetworks have drawn interest. While hypernetworks work well, training them is cumber…