7 papers
RTLScout: Joint Agentic Code and Synthesis Optimization for Efficient Digital Circuits
Felix Arnold, Ryan Amaudruz, Dimitrios Tsaras +2
We present RTLScout, an autonomous system that combines LLM-driven agentic design with circuit-level synthesis optimization and arithmetic architecture sweeps. An LLM agent iterati…
OffQ: Taming Structured Outliers in LLM Quantization by Offsetting
Haoqi Wang, Lorenz K. Mueller, Jiawei Zhuang +2
Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost and memory usage. However, act…
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
Lorenz K. Muller, Philippe Bich, Chiara Boretti +3
Test-time scaling is a powerful approach to obtain better reasoning in large language models, but it becomes memory-bottlenecked during long-horizon decoding, as the KV-cache grows…
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
Luca Benfenati, Matteo Risso, Andrea Vannozzi +5
Key-value (KV) caching enables fast autoregressive decoding but at long contexts becomes a dominant bottleneck in High Bandwidth Memory (HBM) capacity and bandwidth. A common mitig…
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
Ahmet Caner Yüzügüler, Ahmet Ãelik, Jiawei Zhuang +1
Multi-Head Latent Attention (MLA) is a recent attention mechanism adopted in state-of-the-art LLMs such as DeepSeek-v3 and Kimi K2. Thanks to its novel formulation, MLA allows two…
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
Lorenz K. Müller, Philippe Bich, Jiawei Zhuang +3
Post-training quantization has emerged as the most widely used strategy for deploying large language models at low precision. Still, current methods show perplexity degradation at…