4 citations · 4 across the 4 of their papers we have counts for
5 papers · 1 filter
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
Pierre Abillama, Changwoo Lee, Juechu Dong +3
Recent advances in transformer-based foundation models have made them the default choice for many tasks, but their rapidly growing size makes fitting a full model on a single GPU i…
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
Can Yaras, Alec S. Xu, Pierre Abillama +2
Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism. In t…
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference
Changwoo Lee, Soo Min Kwon, Qing Qu +1
Large-scale foundation models have demonstrated exceptional performance in language and vision tasks. However, the numerous dense matrix-vector operations involved in these large n…
Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
Changwoo Lee, Hun-Seok Kim
This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge aris…
A Confidence-Calibrated MOBA Game Winner Predictor
Dong-Hee Kim, Changwoo Lee, Ki-Seok Chung
In this paper, we propose a confidence-calibration method for predicting the winner of a famous multiplayer online battle arena (MOBA) game, League of Legends. In MOBA games, the d…