12 citations · 13 across the 5 of their papers we have counts for
7 papers
CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation
Shuai Tan, Biao Gong, Ke Ma +5
Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi-subject rendering. While existing methods exc…
mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs
Chuan Xu, Qiaosheng Chen, Yutong Feng +1
Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing the capabilities of large language models. However, existing RAG evaluation predominantly focu…
Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan, Ang Wang, Baole Ai +58
This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…
Mimir: Improving Video Diffusion Models for Precise Text Understanding
Shuai Tan, Biao Gong, Yutong Feng +6
Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features…
In-Context LoRA for Diffusion Transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
Recent research arXiv:2410.15027 has explored the use of diffusion transformers (DiTs) for task-agnostic image generation by simply concatenating attention tokens across images. Ho…
Group Diffusion Transformers are Unsupervised Multitask Learners
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
While large language models (LLMs) have revolutionized natural language processing with their task-agnostic capabilities, visual generation tasks such as image translation, style t…