1 paper
Ziyin Xiong, Nikolaos Gkanatsios, Nikos Gkanatsios +2
We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate…