1 paper
Andrew Ting Yan Li, Zhuo Li, Zhelin Yang +3
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; adding causal reasoning improves manipulation, but current methods pay for…