5 papers
Image Generators are Generalist Vision Learners
Valentin Gabeur, Shangbang Long, Songyou Peng +22
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language under…
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
Boyang Deng, Songyou Peng, Kyle Genova +4
We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in t…
SplatTalk: 3D VQA with Gaussian Splatting
Anh Thai, Songyou Peng, Kyle Genova +2
Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3…
3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation
Zihao Xiao, Longlong Jing, Shangxuan Wu +9
3D panoptic segmentation is a challenging perception task, especially in autonomous driving. It aims to predict both semantic and instance annotations for 3D points in a scene. Alt…
Gaussian3Diff: 3D Gaussian Diffusion for 3D Full Head Synthesis and Editing
Yushi Lan, Feitong Tan, Di Qiu +8
We present a novel framework for generating photorealistic 3D human head and subsequently manipulating and reposing them with remarkable flexibility. The proposed approach leverage…