6 papers
A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology
Dichang Zhang, Jiaqi Deng, Yixuan Shao +7
Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology recognition tasks still requir…
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
Nurislam Tursynbek, Zhiqiang Lao, Heather Yu +2
Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering, drifting, or unstable motio…
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
Jiali Cui, Zhiqiang Lao, Heather Yu
Energy-based models (EBMs) are a flexible class of deep generative models and are well-suited to capture complex dependencies in multimodal data. However, learning multimodal EBM b…
OT-Talk: Animating 3D Talking Head with Optimal Transportation
Xinmu Wang, Xiang Gao, Xiyun Song +4
Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech s…
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Sakib Reza, Xiyun Song, Heather Yu +3
Integrating vision models into large language models (LLMs) has sparked significant interest in creating vision-language foundation models, especially for video understanding. Rece…
OccludeNeRF: Geometric-aware 3D Scene Inpainting with Collaborative Score Distillation in NeRF
Jingyu Shi, Achleshwar Luthra, Jiazhi Li +5
With Neural Radiance Fields (NeRFs) arising as a powerful 3D representation, research has investigated its various downstream tasks, including inpainting NeRFs with 2D images. Desp…