1 paper · 1 filter
Zebin You, Shen Nie, Xiaolu Zhang +5
In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, represen…