3 papers
cs.RO2026
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
Fawad Javed Fateh, Ali Shah Ali, Murad Popattia +4
We present a novel hierarchical spatiotemporal action tokenizer for in-context imitation learning. We first propose a hierarchical approach, which consists of two successive levels…
cs.LG2026
Mixing Times of Glauber Dynamics on Masked Language Models
Suvadip Sana, Sami Wolf, Neer Mehta +4
Masked language models (MLMs) define local conditional distributions over tokens but do not, in general, correspond to any consistent joint distribution over sequences. This raises…
cs.CV2025
Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
Syed Ahmed Mahmood, Ali Shah Ali, Umer Ahmed +3
We study self-supervised procedure learning, which discovers key steps and their order from a set of unlabeled videos. Previous methods typically learn frame-to-frame correspondenc…