8 papers
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
Kunpeng Song, Tingbo Hou, Zecheng He +12
In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-…
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
Hongjie Wang, Chih-Yao Ma, Yen-Cheng Liu +10
Text-to-video generation enhances content creation but is highly computationally intensive: The computational cost of Diffusion Transformers (DiTs) scales quadratically in the numb…
FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
Bichen Wu, Chaojian Li, Hang Zhang +6
Neural Architecture Search (NAS) has been widely adopted to design accurate and efficient image classification models. However, applying NAS to a new computer vision task still req…
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
Feng Liang, Haoyu Ma, Zecheng He +10
Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personal…
Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile Telepresence
Yonggan Fu, Yuecheng Li, Chenghui Li +4
Real-time and robust photorealistic avatars for telepresence in AR/VR have been highly desired for enabling immersive photorealistic telepresence. However, there still exists one k…
Towards Automated Model Design on Recommender Systems
Tunhou Zhang, Dehua Cheng, Yuchen He +10
The increasing popularity of deep learning models has created new opportunities for developing AI-based recommender systems. Designing recommender systems using deep neural network…