2 papers
cs.LG2026
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
Wilhelm Tranheden, Shahnawaz Ahmed, Devdatt Dubhashi +2
Language models are increasingly adopting smaller architectures optimized for consumer devices. In this setting, inference efficiency is the primary constraint. Meanwhile, vocabula…
quant-ph2026
Train on classical, deploy on quantum: scaling generative quantum machine learning to a thousand qubits
Erik Recio-Armengol, Shahnawaz Ahmed, Joseph Bowles
We propose an approach to generative quantum machine learning that overcomes the fundamental scaling issues of variational quantum circuits. The core idea is to use a class of gene…