3 citations · 3 across the 2 of their papers we have counts for
10 papers
FlashFace: Human Image Personalization with High-fidelity Identity Preservation
Shilong Zhang, Lianghua Huang, Xi Chen +7
This work presents FlashFace, a practical tool with which users can easily personalize their own photos on the fly by providing one or a few reference face images and a text prompt…
CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation
Shuai Tan, Biao Gong, Ke Ma +5
Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi-subject rendering. While existing methods exc…
mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs
Chuan Xu, Qiaosheng Chen, Yutong Feng +1
Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing the capabilities of large language models. However, existing RAG evaluation predominantly focu…
Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan, Ang Wang, Baole Ai +58
This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…
Mimir: Improving Video Diffusion Models for Precise Text Understanding
Shuai Tan, Biao Gong, Yutong Feng +6
Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features…
In-Context LoRA for Diffusion Transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
Recent research arXiv:2410.15027 has explored the use of diffusion transformers (DiTs) for task-agnostic image generation by simply concatenating attention tokens across images. Ho…