4 papers · 1 filter
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
Tao Lu, Haoyu Wang, Zonghui Wang +3
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can acce…
On the (Generative) Linear Sketching Problem
Xinyu Yuan, Yan Qiao, Zonghui Wang +1
Sketch techniques have been extensively studied in recent years and are especially well-suited to data streaming scenarios, where the sketch summary is updated quickly and compactl…
Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models
Xinyu Yuan, Yan Qiao, Zonghui Wang +1
The multi-commodity flow (MCF) problem is a fundamental topic in network flow and combinatorial optimization, with broad applications in transportation, communication, and logistic…
Learning-based Sketches for Frequency Estimation in Data Streams without Ground Truth
Xinyu Yuan, Yan Qiao, Meng Li +4
Estimating the frequency of items on the high-volume, fast data stream has been extensively studied in many areas, such as database and network measurement. Traditional sketches pr…