2 papers
cs.CV2026
EVEWorld: Physical Evolution Supervision for Embodied World Models
Kaiqi Wang, Songxin Zhang, Zejian Xie +7
Embodied world models enable scalable simulation of embodied interactions for robot learning. However, existing models are prone to Model Laziness, as they focus on visual fidelity…
cs.CV2026
PVI: Plug-in Visual Injection for Vision-Language-Action Models
Zezhou Zhang, Songxin Zhang, Xiao Xiong +8
VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for…