182 citations · 192 across the 7 of their papers we have counts for
6 papers
Guac: Energy-Aware and SSA-Based Generation of Coarse-Grained Merged Accelerators from LLVM-IR
Iulian Brumar, Rodrigo Rocha, Alex Bernat +3
Designing accelerators for resource- and power-constrained applications is a daunting task. High-level Synthesis (HLS) addresses these constraints through resource sharing, an opti…
S: Increasing GPU Utilization during Generative Inference for Higher Throughput
Yunho Jin, Chun-Feng Wu, David Brooks +1
Generating texts with a large language model (LLM) consumes massive amounts of memory. Apart from the already-large model parameters, the key/value (KV) cache that holds informatio…
Design Space Exploration and Optimization for Carbon-Efficient Extended Reality Systems
Mariam Elgamal, Doug Carmean, Elnaz Ansari +8
As computing hardware becomes more specialized, designing environmentally sustainable computing systems requires accounting for both hardware and software parameters. Our goal is t…
MP-Rec: Hardware-Software Co-Design to Enable Multi-Path Recommendation
Samuel Hsia, Udit Gupta, Bilge Acun +5
Deep learning recommendation systems serve personalized content under diverse tail-latency targets and input-query loads. In order to do so, state-of-the-art recommendation models…
PerfSAGE: Generalized Inference Performance Predictor for Arbitrary Deep Learning Models on Edge Devices
Yuji Chai, Devashree Tripathy, Chuteng Zhou +6
The ability to accurately predict deep neural network (DNN) inference performance metrics, such as latency, power, and memory footprint, for an arbitrary DNN on a target hardware p…
Fathom: Reference Workloads for Modern Deep Learning Methods
Robert Adolf, Saketh Rama, Brandon Reagen +2
Deep learning has been popularized by its recent successes on challenging artificial intelligence problems. One of the reasons for its dominance is also an ongoing challenge: the n…