2 papers
cs.AR2025
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
Zhiwen Mo, Lei Wang, Jianyu Wei +8
Large Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency. Such low-bit…
cs.AR2025
CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Qingyuan Liu, Liyan Chen, Yanning Yang +6
Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Ful…