2 papers
cs.AR2026
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
Wenqiang Wang, Yijia Zhang, Zikai Zhang +4
As large language models (LLMs) demonstrate powerful capabilities, deploying them on edge devices has become increasingly crucial, offering advantages in privacy and real-time inte…
cs.AR2026
TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence
Hongyi Guan, Yijia Zhang, Wenqiang Wang +4
The deployment of Large Language Models (LLMs) for real-time intelligence on edge devices is rapidly growing. However, conventional hardware architectures face a fundamental memory…