activity
20192023
most citedLearning Low-Rank Approximation for CNNs

13 citations · 40 across the 15 of their papers we have counts for

collaborators
Showing 2023Show all

8 papers · 1 filter

cs.DC2023

MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems

Samuel Hsia, Alicia Golden, Bilge Acun +5

Training and deploying large-scale machine learning models is time-consuming, requires significant distributed computing infrastructures, and incurs high operational costs. Our ana…

cs.SE2023

Guess & Sketch: Language Model Guided Transpilation

Celine Lee, Abdulrahman Mahmoud, Michal Kurek +5

Maintaining legacy software requires many software and systems engineering hours. Assembly code programs, which demand low-level control over the computer machine state and have no…

cs.CL2023★ 4 cited

INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank Adaptation

Yuji Chai, John Gkountouras, Glenn G. Ko +2

We introduce a method that dramatically reduces fine-tuning VRAM requirements and rectifies quantization errors in quantized Large Language Models. First, we develop an extremely m…

cs.AR2023★ 7 cited

S: Increasing GPU Utilization during Generative Inference for Higher Throughput

Yunho Jin, Chun-Feng Wu, David Brooks +1

Generating texts with a large language model (LLM) consumes massive amounts of memory. Apart from the already-large model parameters, the key/value (KV) cache that holds informatio…

cs.AR2023★ 2 cited

Design Space Exploration and Optimization for Carbon-Efficient Extended Reality Systems

Mariam Elgamal, Doug Carmean, Elnaz Ansari +8

As computing hardware becomes more specialized, designing environmentally sustainable computing systems requires accounting for both hardware and software parameters. Our goal is t…

cs.AR2023

MP-Rec: Hardware-Software Co-Design to Enable Multi-Path Recommendation

Samuel Hsia, Udit Gupta, Bilge Acun +5

Deep learning recommendation systems serve personalized content under diverse tail-latency targets and input-query loads. In order to do so, state-of-the-art recommendation models…