collaborators

7 papers

cs.LG2026

AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs

Gunho Park, Jeongin Bae, Beomseok Kwon +3

The deployment of large language models (LLMs) is increasingly constrained by memory and latency bottlenecks, motivating the need for quantization techniques that flexibly balance…

cs.LG2025

CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs

Gunho Park, Jeongin Bae, Byeongwook Kim +5

Weight-only quantization is widely used to mitigate the memory-bound nature of LLM inference. Codebook-based methods extend this trend by achieving strong accuracy in the extremely…

cs.LG2025

An Inquiry into Datacenter TCO for LLM Inference with FP8

Jiwoo Kim, Joonhyung Lee, Gunho Park +4

As large language models (LLMs) continue to scale, the high power consumption of AI accelerators in datacenters presents significant challenges, substantially increasing the total…

cs.CL2025

Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models

Seungcheol Park, Jeongin Bae, Beomseok Kwon +5

How can we quantize large language models while preserving accuracy? Quantization is essential for deploying large language models (LLMs) efficiently. Binary-coding quantization (B…

cs.LG2025

To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability

Joonhyung Lee, Jeongin Bae, Byeongwook Kim +2

The massive computational costs associated with large language model (LLM) pretraining have spurred great interest in reduced-precision floating-point representations to accelerate…

cs.AR2025

Faster Inference of LLMs using FP8 on the Intel Gaudi

Joonhyung Lee, Shmulik Markovich-Golan, Daniel Ohayon +9

Low-precision data types are essential in modern neural networks during both training and inference as they enhance throughput and computational capacity by better exploiting avail…