1 paper
Zhewen Wan, Tianchen Song, Chen Lin +2
Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional at…