2 papers
cs.CV2026
PVI: Plug-in Visual Injection for Vision-Language-Action Models
Zezhou Zhang, Songxin Zhang, Xiao Xiong +8
VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for…
cs.CV2026
A Prediction-as-Perception Framework for 3D Object Detection
Song Zhang, Haoyu Chen, Ruibo Wang
Humans combine prediction and perception to observe the world. When faced with rapidly moving birds or insects, we can only perceive them clearly by predicting their next position…