agent harness design 1long video processing 1multimodal information extraction 1multimodal retrieval 1provenance tracking 1scientific curation 1sparse token selection 1structured data extraction 1vision-language models 1visual retrieval 1
From the 2 of 20 linked papers with an AI index.
Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
Spatially Grounded Long-Horizon Task Planning in the Wild
Sehun Jung, HyunJee Song, Dong-Hee Kim +4
Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action pla…
cs.RO2025
Latent Action Pretraining from Videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon +13
We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot actio…