7 papers
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
Hongyaoxing Gu, Xinzhe Chen, Lijuan Hu +1
Mixture-of-Experts (MoE) models achieve remarkable performance by sparsely activating specialized experts, yet their massive parameters in experts pose significant challenges for d…
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
Ruijie Zhang, Haozhe Liang, Da Chang +4
Long-context LLM inference is bottlenecked by the memory and bandwidth cost of reading large KV caches during decoding. KV compression reduces this cost by keeping only part of the…
Millikelvin digital-to-analog converter for superconducting quantum processors
Ruizi Hu, Zongyuan Li, Zhancheng Yao +19
Scaling superconducting quantum processors is increasingly constrained by the wiring, heat load, and calibration overhead associated with delivering high-resolution analog signals…
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
Hongyaoxing Gu, Lijuan Hu, Liye Yu +2
Post-training quantization (PTQ) enables effective model compression while preserving relatively high accuracy. Current weight-only PTQ methods primarily focus on the challenging s…
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
Hongyaoxing Gul, Lijuan Hu, Shuzi Niu +1
Traditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, exist…
Identify and Quantify Various Dissipation Mechanisms of Josephson Junction in Superconducting Circuits
Hao Deng, Huijuan Zhan, Lijuan Hu +12
Pinpointing the dissipation mechanisms and evaluating their impacts to the performance of Josephson junction (JJ) are crucial for its application in superconducting circuits. In th…