3d representation 1cross-modal learning 1depth perception 1gaussian primitives 1multimodal fusion 1representation learning 1robotic manipulation 1semantic grounding 1unsupervised semantic segmentation 1vision-language-action 1
From the 2 of 3 linked papers with an AI index.
3 papers
cs.RO2026
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
Mohan Liu, Zhihao Gu, Xuanyu Chen +5
VistaVLA is a two-stage framework that builds a geometry- and semantic-aware 3D cognitive map using Gaussian primitives and compresses it into compact tokens for vision‑language‑ac…
cs.CV2026
UMSS: Towards Unsupervised Multi-modal Semantic Segmentation
Haitian Zhang, Thai Duy Nguyen, Xiangyuan Wang +2
The paper introduces UniM2, an unsupervised framework for multimodal semantic segmentation that learns a shared latent space across sensors using cross‑modal correspondence and a h…
cs.RO2026
Physics-Guided Biomechanical Gait Adaptation for Humanoid Locomotion on Extreme Sloped Terrains
Xuanyu Chen, Mohan Liu, Dengchen Mei +6
Model-free reinforcement learning has enabled impressive humanoid locomotion; however, control on steep slopes remains largely unexplored. Unlike flat or discrete terrains, sloped…