collaborators

6 papers

cs.LG2026

FPTQuant: Function-Preserving Transforms for LLM Quantization

Boris van Breugel, Yelysei Bondarenko, Paul Whatmough +1

Large language models (LLMs) require substantial compute, and thus energy, at inference time. While quantizing weights and activations is effective at improving efficiency, naive q…

cs.CV2026

MobileWan: Closing the Quality Gap for Mobile Video Diffusion

Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv +9

Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coheren…

cs.LG2026

Dissecting Quantization Error: A Concentration-Alignment Perspective

Marco Federici, Boris van Breugel, Paul Whatmough +1

Quantization can drastically increase the efficiency of large language and vision models, but typically incurs an accuracy drop. Recently, function-preserving transforms (e.g. rota…

cs.LG2025

STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization

Marco Federici, Riccardo Del Chiaro, Boris van Breugel +2

Quantization is the key method for reducing inference latency, power and memory footprint of generative AI models. However, accuracy often degrades sharply when activations are qua…

cs.CV2025

HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations

Marco Federici, Riccardo Del Chiaro, Boris van Breugel +2

Diffusion models represent the cutting edge in image generation, but their high memory and computational demands hinder deployment on resource-constrained devices. Post-Training Qu…

cs.LG2025

Position: All Current Generative Fidelity and Diversity Metrics are Flawed

Ossi Räisä, Boris van Breugel, Mihaela van der Schaar

Any method's development and practical application is limited by our ability to measure its reliability. The popularity of generative modeling emphasizes the importance of good syn…