energy efficiency 1fast hadamard transform 1hardware accelerator 1large language models 1low-bit quantization 1
From the 1 of 4 linked papers with an AI index.
6 citations · 6 across the 3 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026★ 6 cited
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
Sangjin Kim, Yuseon Choi, Jungjun Oh +2
LightRot introduces a lightweight rotation scheme and a dedicated hardware accelerator that enable energy‑efficient, low‑bit inference for large language models such as LLaMA2‑13B…
cs.AR2026
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
Yuseon Choi, Sangjin Kim, Jungjun Oh +3
MoE models offer efficient scaling through conditional computation, but their large parameter size and expensive expert offloading make on-device deployment challenging. Existing a…