From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
Hyunwoo Oh, Suyeon Jang, Hanning Chen +4
PolyQ is a co-designed compiler and quantization framework that assigns per‑channel bit‑widths to LLM activations on CPUs, enabling fine‑grained fractional‑bit precision while keep…
cs.AR2025
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
Hyunwoo Oh, KyungIn Nam, Rajat Bhattacharjya +7
Recent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. W…