1 citations · 2 across the 7 of their papers we have counts for
8 papers
InnoText: A Unified Model for Visual Text Generation and Editing
Haowei Liu, Runze He, Jian Lu +10
Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing remains relatively underexp…
Graph of States: Solving Abductive Tasks with Large Language Models
Yu Luo, Rongchen Gao, Lu Teng +9
Logical reasoning encompasses deduction, induction, and abduction. However, while Large Language Models (LLMs) have effectively mastered the former two, abductive reasoning remains…
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma, Jiasong Feng, Ke Cao +4
Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue…
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
Jing Wang, Ao Ma, Ke Cao +9
Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle…
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
Ke Cao, Jing Wang, Ao Ma +11
The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled di…
Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task
Jing Wang, Ao Ma, Jiasong Feng +3
The global self-attention mechanism in diffusion transformers involves redundant computation due to the sparse and redundant nature of visual information, and the attention map of…