2 papers
cs.RO2026
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
Jingxian Lu, Wenke Xia, Yuxuan Wu +2
Proprioceptive information is critical for precise servo control by providing real-time robotic states. Its collaboration with vision is highly expected to enhance performances of…
cs.LG2025
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Zebin You, Shen Nie, Xiaolu Zhang +5
In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, represen…