embodied control 1latency reduction 1predictive representations 1visual-language grounding 1world models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
Fulong Liu, Liang Xu, Chengqun Yang +3
Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distr…
cs.RO2026
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Weili Zeng, Yitong Xing, Fulong Liu +10
The paper introduces Enfold, a method that folds the computation of a world-generative model into a predictive representation derived from the current visual scene and language ins…