2 citations · 3 across the 12 of their papers we have counts for
Showing 2024Show all
3 papers · 1 filter
cs.CV2024★ 1 cited
EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
Jianrong Zhang, Hehe Fan, Yi Yang
Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diff…
cs.CV2024★ 2 cited
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
Yue Zhang, Hehe Fan, Yi Yang
To bridge the gap between vision and language modalities, Multimodal Large Language Models (MLLMs) usually learn an adapter that converts visual inputs to understandable tokens for…
cs.CV2024
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
Wei Li, Hehe Fan, Yongkang Wong +2
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability…