2 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.RO2026
Mask World Model: Predicting What Matters for Robust Robot Policy Learning
Yunfan Lou, Xiaowei Chi, Xiaojie Zhang +9
World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often fo…
cs.CV2025
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
Chenghao Liu, Jiachen Zhang, Chengxuan Li +4
Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This fram…
cs.CV2024★ 2 cited
A Survey on Long Video Generation: Challenges, Methods, and Prospects
Chengxuan Li, Di Huang, Zeyu Lu +3
Video generation is a rapidly advancing research area, garnering significant attention due to its broad range of applications. One critical aspect of this field is the generation o…