6 citations · 8 across the 3 of their papers we have counts for
Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026★ 6 cited
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
Sangjin Kim, Yuseon Choi, Jungjun Oh +2
As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes…
cs.AR2026★ 2 cited
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
Sangjin Kim, Yuseon Choi, Byeongcheol Kim +2
Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often…
cs.AR2025
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
Yuseon Choi, Sangjin Kim, Jungjun Oh +3
MoE models offer efficient scaling through conditional computation, but their large parameter size and expensive expert offloading make on-device deployment challenging. Existing a…