2 papers
cs.LG2025
Squat: Quant Small Language Models on the Edge
Xuan Shen, Peiyan Dong, Zhenglun Kong +9
A growing trend has emerged in designing high-quality Small Language Models (SLMs) with a few million parameters. This trend is driven by the increasing concerns over cloud costs,…
cs.AR2025
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
Zhaoyang Han, Yicheng Qian, Michael Zink +1
Sketch-based algorithms for network traffic monitoring have drawn increasing interest in recent years due to their sub-linear memory efficiency and high accuracy. As the volume of…