6 papers
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Qi Li, Yanzhe Zhao, Yongxin Zhou +4
Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various modalities for a given query. Ho…
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
Yuan Zhou, Shilong Jin, Litao Hua +3
Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods le…
Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
Shilong Jin, Haoran Duan, Litao Hua +2
Versatile 3D tasks (e.g., generation or editing) that distill from Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D…
ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
Yuan Zhou, Litao Hua, Shilong Jin +2
Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information acr…
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Shixiong Xu, Chenghao Zhang, Lubin Fan +5
Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained s…
Attention in Diffusion Model: A Survey
Litao Hua, Fan Liu, Jie Su +9
Attention mechanisms have become a foundational component in diffusion models, significantly influencing their capacity across a wide range of generative and discriminative tasks.…