2 papers
cs.CV2026
Coverage-Driven Adaptive Keyframe Selection for Video Understanding
Junyang Zhang, Puhan Luo, Chen Tang +2
Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substan…
cs.RO2026
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
Yushan Liu, Peibo Sun, Xintao Chao +8
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level v…