2 papers
cs.RO2026
FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference
Zekai Li, Jiaming Tang, Zhijian Liu
Vision-Language-Action (VLA) models are increasingly promising for robotic manipulation, yet their real-world deployment remains bottlenecked by high inference latency and unstable…
cs.CV2026
TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling
Ruohan Wu, Ziqi Zhu, Yang Zhao +6
Training high-resolution AI-based Earth forecasting models is memory-intensive. Window-based Swin Transformers reduce the quadratic cost of global attention, but existing distribut…