2 papers
cs.RO2026
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
Ruisen Tu, Arth Shukla, Sohyun Yoo +5
Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning…
cs.CV2026
PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
Xiang Zhang, Sohyun Yoo, Hongrui Wu +3
We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed…