6 papers
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
Qunyou Liu, Marina Zapater, David Atienza
Transformers have revolutionized AI in natural language processing and computer vision, but their large computation and memory demands pose major challenges for hardware accelerati…
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
Qunyou Liu, Pengbo Yu, Marina Zapater +1
Deep neural networks (DNNs) are essential for performing advanced tasks on edge or mobile devices, yet their deployment is often hindered by severe resource constraints, including…
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
Po-Kai Hsu, Weihong Xu, Qunyou Liu +2
Retrieval-Augmented Generation (RAG) relies on large-scale Approximate Nearest Neighbor Search (ANNS) to retrieve semantically relevant context for large language models. Among ANN…
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
Qunyou Liu, Darong Huang, Marina Zapater +1
Large Language Models (LLMs) are becoming the backbone of modern cloud services, yet their inference costs are dominated by GPU energy. Unlike traditional GPU workloads, LLM infere…
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
Qunyou Liu, Marina Zapater, David Atienza
Transformers are central to advances in artificial intelligence (AI), excelling in fields ranging from computer vision to natural language processing. Despite their success, their…
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
Qunyou Liu, Marina Zapater, David Atienza
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelera…