3 papers
cs.LG2026
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models
Yuval Domb, Hadar Sackstein, Tomer Solberg
We present HyperQuant (Hadamard, optimallY Packing, Entropy Rice-coding), a unified post-training quantization pipeline for the weights and the KV cache of large language and diffu…
cs.IT2026
Why Self-Supervised Encoders Want to Be Normal
Yuval Domb
Self-supervised learning has achieved remarkable empirical success in learning robust representations without explicit labels, most recently demonstrated within the framework of Jo…
cs.CV2025
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
Dor Shmilovich, Tony Wu, Aviad Dahan +1
Diffusion Transformers, particularly for video generation, achieve remarkable quality but suffer from quadratic attention complexity, leading to prohibitive latency. Existing accel…