4 papers
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi +3
Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually expand their capabilities, ma…
FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
Jiacheng Liu, Peiliang Cai, Qinming Zhou +9
The application of diffusion transformers is suffering from their significant inference costs. Recently, feature caching has been proposed to solve this problem by reusing features…
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
Yi Chen, Sen Liang, Zixiang Zhou +6
Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving chara…
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
Ruihuang Li, Caijin Zhou, Shoujian Zheng +55
Intelligent game creation represents a transformative advancement in game development, utilizing generative artificial intelligence to dynamically generate and enhance game content…