6 papers
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
Bakbergen Ryskulov, Iker García-Ferrero, David Montero +5
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together t…
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Bakbergen Ryskulov, Iker García-Ferrero, Iker GarcÃa-Ferrero +7
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model…
Room-temperature tuning and probing of Fermi polarons in atomically thin semiconductors on a plasmonic metasurface
Tingting Wu, Francesca Maria Marchetti, Antonio Tiene +9
The Fermi polaron, arising from interactions between a mobile impurity and a degenerate Fermi sea, is a many-body quasiparticle that provides a sensitive probe of strongly correlat…
Scaling Laws for Energy Efficiency of Local LLMs
Ander Alvarez, Alessandro Genuardi, Nilotpal Sinha +6
Deploying local large language models and vision-language models on edge devices requires balancing accuracy with constrained computational and energy budgets. Although graphics pr…
Efficient calculation of trion energies in monolayer transition metal dichalcogenides
Sangeet S. Kumar, Brendan C. Mulkerin, Antonio Tiene +3
The reduced dielectric screening in atomically thin semiconductors leads to remarkably strong electron interactions. As a result, bound electron-hole pairs (excitons) and charged e…
Exact Quantum Virial Expansion for the Optical Response of Doped Two-Dimensional Semiconductors
B. C. Mulkerin, A. Tiene, F. M. Marchetti +2
We present a quantum virial expansion for the optical response of a doped two-dimensional semiconductor. As we show, this constitutes a perturbatively exact theory in the high-temp…