Publications (14)
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats
Giuseppe Franco, Ian Colbert, Pablo Monteagudo-Lago +2
The paper presents dMX, a differentiable framework that learns per-layer floating‑point bit‑widths for large language models, enabling mixed‑precision quantization that balances ac…
Memory-Efficient Dataflow Inference for Deep CNNs on FPGA
Lucian Petrica, Tobias Alonso, Mairin Kroes +3
Custom dataflow Convolutional Neural Network (CNN) inference accelerators on FPGA are tailored to a specific CNN topology and store parameters in On-Chip Memory (OCM), resulting in…
From closed to open access: A case study of flipped journals
Fakhri Momeni, Nicholas Fraser, Isabella Peters +1
In recent years, increased stakeholder pressure to transition research to Open Access has led to many journals "flipping" from a toll access to an open access publishing model. Cha…
Improving Quantization with Post-Training Model Expansion
Giuseppe Franco, Pablo Monteagudo-Lago, Ian Colbert +2
The size of a model has been a strong predictor of its quality, as well as its cost. As such, the trade-off between model cost and quality has been well-studied. Post-training opti…
Compressing Low Precision Deep Neural Networks Using Sparsity-Induced Regularization in Ternary Networks
Julian Faraone, Nicholas Fraser, Giulio Gambardella +2
A low precision deep neural network training technique for producing sparse, ternary neural networks is presented. The technique incorporates hard- ware implementation costs during…
Optimal Post-Training Quantization Scales and Where to Find Them
Juan Amboage, Pablo Monteagudo-Lago, Ian Colbert +2
Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations. The scaling factor that defines the quantization grid is typically…
What happens when a journal converts to Open Access? A bibliometric analysis
Fakhri Momeni, Philipp Mayr, Nicholas Fraser +1
In recent years, increased stakeholder pressure to transition research to Open Access has led to many journals converting, or 'flipping', from a closed access (CA) to an open acces…
SYQ: Learning Symmetric Quantization For Efficient Deep Neural Networks
Julian Faraone, Nicholas Fraser, Michaela Blott +1
Inference for state-of-the-art deep neural networks is computationally expensive, making them difficult to deploy on constrained hardware environments. An efficient way to reduce t…
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
Jiahao Zhang, Zifan He, Nicholas Fraser +3
We present FlexLLM, a composable High-Level Synthesis (HLS) library for rapid development of domain-specific LLM accelerators. FlexLLM exposes key architectural degrees of freedom…
No Deal: Investigating the Influence of Restricted Access to Elsevier Journals on German Researchers' Publishing and Citing Behaviours
Nicholas Fraser, Anne Hobert, Najko Jahn +2
In 2014, a union of German research organisations established Projekt DEAL, a national-level project to negotiate licensing agreements with large scientific publishers. Negotiation…
Scaling Neural Network Performance through Customized Hardware Architectures on Reconfigurable Logic
Michaela Blott, Thomas B. Preusser, Nicholas Fraser +4
Convolutional Neural Networks have dramatically improved in recent years, surpassing human accuracy on certain problems and performance exceeding that of traditional computer visio…
Inference of Quantized Neural Networks on Heterogeneous All-Programmable Devices
Thomas B. PreuÃer, Giulio Gambardella, Nicholas Fraser +1
Neural networks have established as a generic and powerful means to approach challenging problems such as image classification, object detection or decision making. Their successfu…
FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
Michaela Blott, Thomas Preusser, Nicholas Fraser +3
Convolutional Neural Networks have rapidly become the most successful machine learning algorithm, enabling ubiquitous machine vision and intelligent decisions on even embedded comp…
Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
Sean O. Settle, Manasa Bollavaram, Paolo D'Alberto +6
Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly…