collaborators

6 papers

cs.CV2026

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

Qi Li, Yanzhe Zhao, Yongxin Zhou +4

Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various modalities for a given query. Ho…

cs.CV2026

ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation

Yuan Zhou, Shilong Jin, Litao Hua +3

Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods le…

cs.CV2025

Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting

Shilong Jin, Haoran Duan, Litao Hua +2

Versatile 3D tasks (e.g., generation or editing) that distill from Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D…

cs.CV2025

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

Yuan Zhou, Litao Hua, Shilong Jin +2

Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information acr…

cs.CV2025

AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

Shixiong Xu, Chenghao Zhang, Lubin Fan +5

Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained s…

cs.LG2025

Attention in Diffusion Model: A Survey

Litao Hua, Fan Liu, Jie Su +9

Attention mechanisms have become a foundational component in diffusion models, significantly influencing their capacity across a wide range of generative and discriminative tasks.…