5 papers
OT-Talk: Animating 3D Talking Head with Optimal Transportation
Xinmu Wang, Xiang Gao, Xiyun Song +4
Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech s…
Scene Perceived Image Perceptual Score (SPIPS): combining global and local perception for image quality assessment
Zhiqiang Lao, Heather Yu
The rapid advancement of artificial intelligence and widespread use of smartphones have resulted in an exponential growth of image data, both real (camera-captured) and virtual (AI…
ePBR: Extended PBR Materials in Image Synthesis
Yu Guo, Zhiqiang Lao, Xiyun Song +3
Realistic indoor or outdoor image synthesis is a core challenge in computer vision and graphics. The learning-based approach is easy to use but lacks physical consistency, while tr…
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Sakib Reza, Xiyun Song, Heather Yu +3
Integrating vision models into large language models (LLMs) has sparked significant interest in creating vision-language foundation models, especially for video understanding. Rece…
OccludeNeRF: Geometric-aware 3D Scene Inpainting with Collaborative Score Distillation in NeRF
Jingyu Shi, Achleshwar Luthra, Jiazhi Li +5
With Neural Radiance Fields (NeRFs) arising as a powerful 3D representation, research has investigated its various downstream tasks, including inpainting NeRFs with 2D images. Desp…