2 papers
cs.RO2026
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
Jiahui Niu, Kefan Gu, Yucheng Zhao +5
Diffusion-based vision-language-action models (dVLAs) are promising for embodied intelligence but are fundamentally limited in real-time deployment by the high latency of full infe…
cs.AR2026
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
Jinxin Yu, Yudong Pan, Mengdi Wang +4
Transformer-based models dominate modern AI workloads but exacerbate memory bottlenecks due to their quadratic attention complexity and ever-growing model sizes. Existing accelerat…