3 papers
cs.RO2026
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang +11
Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world…
cs.CV2026
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
Zhejia Cai, Yandan Yang, Xinyuan Chang +5
Latent Action Models (LAMs) enable Vision- Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottleneck…
cs.CV2026
Improving Multi-View Reconstruction via Texture-Guided Gaussian-Mesh Joint Optimization
Zhejia Cai, Puhua Jiang, Shiwei Mao +2
Reconstructing real-world objects from multi-view images is essential for applications in 3D editing, AR/VR, and digital content creation. Existing methods typically prioritize eit…