papers

Publications (14)

cs.LG2026

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats

Giuseppe Franco, Ian Colbert, Pablo Monteagudo-Lago +2

The paper presents dMX, a differentiable framework that learns per-layer floating‑point bit‑widths for large language models, enabling mixed‑precision quantization that balances ac…

#mixed-precision quantization#floating-point formats#large language models#model compression
cs.AR2020

Memory-Efficient Dataflow Inference for Deep CNNs on FPGA

Lucian Petrica, Tobias Alonso, Mairin Kroes +3

Custom dataflow Convolutional Neural Network (CNN) inference accelerators on FPGA are tailored to a specific CNN topology and store parameters in On-Chip Memory (OCM), resulting in…

cs.DL2019

From closed to open access: A case study of flipped journals

Fakhri Momeni, Nicholas Fraser, Isabella Peters +1

In recent years, increased stakeholder pressure to transition research to Open Access has led to many journals "flipping" from a toll access to an open access publishing model. Cha…

cs.LG2025

Improving Quantization with Post-Training Model Expansion

Giuseppe Franco, Pablo Monteagudo-Lago, Ian Colbert +2

The size of a model has been a strong predictor of its quality, as well as its cost. As such, the trade-off between model cost and quality has been well-studied. Post-training opti…

cs.CV2017

Compressing Low Precision Deep Neural Networks Using Sparsity-Induced Regularization in Ternary Networks

Julian Faraone, Nicholas Fraser, Giulio Gambardella +2

A low precision deep neural network training technique for producing sparse, ternary neural networks is presented. The technique incorporates hard- ware implementation costs during…

cs.LG2026

Optimal Post-Training Quantization Scales and Where to Find Them

Juan Amboage, Pablo Monteagudo-Lago, Ian Colbert +2

Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations. The scaling factor that defines the quantization grid is typically…

cs.DL2021

What happens when a journal converts to Open Access? A bibliometric analysis

Fakhri Momeni, Philipp Mayr, Nicholas Fraser +1

In recent years, increased stakeholder pressure to transition research to Open Access has led to many journals converting, or 'flipping', from a closed access (CA) to an open acces…

cs.CV2018

SYQ: Learning Symmetric Quantization For Efficient Deep Neural Networks

Julian Faraone, Nicholas Fraser, Michaela Blott +1

Inference for state-of-the-art deep neural networks is computationally expensive, making them difficult to deploy on constrained hardware environments. An efficient way to reduce t…

cs.AR2026

FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design

Jiahao Zhang, Zifan He, Nicholas Fraser +3

We present FlexLLM, a composable High-Level Synthesis (HLS) library for rapid development of domain-specific LLM accelerators. FlexLLM exposes key architectural degrees of freedom…

cs.DL2021

No Deal: Investigating the Influence of Restricted Access to Elsevier Journals on German Researchers' Publishing and Citing Behaviours

Nicholas Fraser, Anne Hobert, Najko Jahn +2

In 2014, a union of German research organisations established Projekt DEAL, a national-level project to negotiate licensing agreements with large scientific publishers. Negotiation…

cs.CV2018

Scaling Neural Network Performance through Customized Hardware Architectures on Reconfigurable Logic

Michaela Blott, Thomas B. Preusser, Nicholas Fraser +4

Convolutional Neural Networks have dramatically improved in recent years, surpassing human accuracy on certain problems and performance exceeding that of traditional computer visio…

cs.NE2018

Inference of Quantized Neural Networks on Heterogeneous All-Programmable Devices

Thomas B. Preußer, Giulio Gambardella, Nicholas Fraser +1

Neural networks have established as a generic and powerful means to approach challenging problems such as image classification, object detection or decision making. Their successfu…

cs.AR2018

FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks

Michaela Blott, Thomas Preusser, Nicholas Fraser +3

Convolutional Neural Networks have rapidly become the most successful machine learning algorithm, enabling ubiquitous machine vision and intelligent decisions on even embedded comp…

cs.LG2018

Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines

Sean O. Settle, Manasa Bollavaram, Paolo D'Alberto +6

Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly…