2 papers
cs.RO2026
CAPE: Contrastive Action-conditioned Parallel Encoding for Embodied Planning
Cong Chen, Haowen Wang, Zhixiang Zhang +2
Embodied agents need to predict the future consequences of candidate actions in order to plan effectively before execution. Existing visual dynamics models learn by reconstructing…
cs.CV2026
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
Lianying Chao, Linfeng Yin, Peiyu Ren +8
Video captioning models convert frames into visual tokens and generate descriptions with large language models (LLMs). Since encoding all frames is prohibitively expensive, uniform…