causal pretraining 1foundation models 1robot control 1sparse mixture of experts 1video-action models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.RO2026
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
cs.CV2026
Learning Probabilistic Embeddings for Unsupervised Action Segmentation
Shuai Li, Duc Manh Vu, Juergen Gall
This paper concerns the problem of unsupervised temporal action segmentation for long, untrimmed videos. Recent successful approaches follow a joint representation learning and clu…