2 papers
cs.CV2026
PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding
Siru Zhong, Qiongyan Wang, Xiaohui Lv +5
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value…
cs.RO2026
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
Lin Liu, Zhicheng Bao, Lu Zhang +7
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100…