2 papers
cs.RO2026
DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
Fenghao Lei, Zhixiong Huang, Long Yang +5
World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generati…
cs.CV2026
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
Lyuke Wang, Zhuo Li, Guangxu Zhu
While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive…