3 papers
cs.RO2026
KEEP: A KV-Cache-Centric Memory Management System for Efficient Embodied Planning
Zebin Yang, Tong Xie, Baotong Lu +3
Memory-augmented Large Language Models (LLMs) have demonstrated remarkable capability for complex and long-horizon embodied planning. By keeping track of past experiences and envir…
cs.LG2024
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
Zebin Yang, Renze Chen, Taiqiang Wu +5
In this paper, we propose MCUBERT to enable language models like BERT on tiny microcontroller units (MCUs) through network and scheduling co-optimization. We observe the embedding…
cs.AR2024
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
Size Zheng, Renze Chen, Meng Li +3
IoT devices based on microcontroller units (MCU) provide ultra-low power consumption and ubiquitous computation for near-sensor deep learning models (DNN). However, the memory of M…