3 papers
cs.LG2026
Float8@2bits: Entropy Coding Enables Data-Free Model Compression
Patrick Putzky, Martin Genzel, Mattes Mollenhauer +3
Post-training compression is currently divided into two contrasting regimes. On the one hand, fast, data-free, and model-agnostic methods (e.g., NF4 or HQQ) offer maximum accessibi…
math.ST2026
Asymptotic e-processes
Pierre-François Massiani, Sebastian Schulze, Mattes Mollenhauer
We investigate the concept of an asymptotic e-process, which is a doubly-indexed stochastic process that possesses, asymptotically for an approximati…
cs.LG2025
Choose Your Model Size: Any Compression of Large Language Models Without Re-Computation
Martin Genzel, Patrick Putzky, Pengfei Zhao +5
The adoption of Foundation Models in resource-constrained environments remains challenging due to their large size and inference costs. A promising way to overcome these limitation…