2 papers
cs.AI2026
From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization
Achille Jacquemond, Yuma Ichikawa, Akira Sakai
Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transfor…
cs.LG2026
OneComp: One-Line Revolution for Generative AI Model Compression
Yuma Ichikawa, Keiji Kimura, Akihiro Yoshida +11
Deploying foundation models is increasingly constrained by memory footprint, latency, and hardware costs. Post-training compression can mitigate these bottlenecks by reducing the p…