1 citations · 1 across the 12 of their papers we have counts for
19 papers · 1 filter
DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
Binglei Li, Mengping Yang, Zhiyu Tan +4
Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of…
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Binglei Li, Mengping Yang, Zhiyu Tan +2
Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and Qwen Image, have facilitated thr…
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
Mengping Yang, Zhiyu Tan, Binglei Li +3
Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturi…
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
Hao Yang, Zhiyu Tan, Jia Gong +7
We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified vi…
Omni-Video: Democratizing Unified Video Understanding and Generation
Zhiyu Tan, Hao Yang, Luozheng Qin +3
Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current fo…
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
Hao Li, Yuhao Wang, Wenning Hao +3
RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing…