2 papers
cs.LG2026
Finer is Better (with the Right Scaling)
Clemens Schaefer, Gil Tabak
Microscaling is a critical technique for preserving the quality of Large Language Models (LLMs) quantized to ultra-low precision formats. Intuitively, finer block sizes should yiel…
cs.LG2025
EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration
Ibrahim Ahmed, Clemens Schaefer, Gil Tabak +5
While Large Language Models (LLMs) have become highly influential, their enormous scale presents significant deployment challenges. Efficiently serving these models typically requi…