collaborators

5 papers

cs.LG2026

SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference

Qunyou Liu, Pengbo Yu, Marina Zapater +1

Deep neural networks (DNNs) are essential for performing advanced tasks on edge or mobile devices, yet their deployment is often hindered by severe resource constraints, including…

cs.PF2025

GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving

Qunyou Liu, Darong Huang, Marina Zapater +1

Large Language Models (LLMs) are becoming the backbone of modern cloud services, yet their inference costs are dominated by GPU energy. Unlike traditional GPU workloads, LLM infere…

cs.AR2025

Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators

Qunyou Liu, Marina Zapater, David Atienza

The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelera…

cs.ET2025

LionHeart: A Layer-based Mapping Framework for Heterogeneous Systems with Analog In-Memory Computing Tiles

Corey Lammie, Yuxuan Wang, Flavio Ponzina +7

When arranged in a crossbar configuration, resistive memory devices can be used to execute Matrix-Vector Multiplications (MVMs), the most dominant operation of many Machine Learnin…

cs.AR2025

MatrixFlow: System-Accelerator co-design for high-performance transformer applications

Qunyou Liu, Marina Zapater, David Atienza

Transformers are central to advances in artificial intelligence (AI), excelling in fields ranging from computer vision to natural language processing. Despite their success, their…