Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs
Ziang Duan, Jiajun Wu, Zetian Chen +8
Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separa…
cs.AR2024
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
Jiajun Wu, Mo Song, Jingmin Zhao +3
Modern transformer-based deep neural networks present unique technical challenges for effective acceleration in real-world applications. Apart from the vast amount of linear operat…