3 citations · 5 across the 9 of their papers we have counts for
5 papers · 1 filter
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
Zebin You, Xiaolu Zhang, Jun Zhou +2
We present \textbf{LLaDA-o}, an effective and length-adaptive omni diffusion model for multimodal understanding and generation. LLaDA-o is built on a Mixture of Diffusion (MoD) fra…
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
Ying Ba, Tianyu Zhang, Yalong Bai +4
Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have fai…
Uniform Attention Maps: Boosting Image Fidelity in Reconstruction and Editing
Wenyi Mo, Tianyu Zhang, Yalong Bai +2
Text-guided image generation and editing using diffusion models have achieved remarkable advancements. Among these, tuning-free methods have gained attention for their ability to p…
Dynamic Prompt Optimizing for Text-to-Image Generation
Wenyi Mo, Tianyu Zhang, Yalong Bai +3
Text-to-image generative models, specifically those based on diffusion models like Imagen and Stable Diffusion, have made substantial advancements. Recently, there has been a surge…
Supporting Vision-Language Model Inference with Confounder-pruning Knowledge Prompt
Jiangmeng Li, Wenyi Mo, Wenwen Qiang +4
Vision-language models are pre-trained by aligning image-text pairs in a common space to deal with open-set visual concepts. To boost the transferability of the pre-trained models,…