9 papers
Scaling Properties of Continuous Diffusion Spoken Language Models
Jason Ramapuram, Eeshan Gunesh Dhekane, Amitis Shidani +6
Speech-only spoken language models (SLMs) lag behind text and text-speech models in performance, with recent discrete autoregressive (AR) SLMs indicating significant computational…
Enabling Differentially Private Federated Learning for Speech Recognition: Benchmarks, Adaptive Optimizers and Gradient Clipping
Martin Pelikan, Sheikh Shams Azam, Vitaly Feldman +4
While federated learning (FL) and differential privacy (DP) have been extensively studied, their application to automatic speech recognition (ASR) remains largely unexplored due to…
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Hyung Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko +5
We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing dist…
Evidence for CP violation in decays
LHCb collaboration, R. Aaij, B. Adeva +698
Three-body and decays are studied using a data sample corresponding to an integrated luminosity of 3.0 collected by…
Measurements of charm mixing and violation using decays
LHCb collaboration, R. Aaij, B. Adeva +766
Measurements of charm mixing and violation parameters from the decay-time-dependent ratio of to decay rates and the charge-conjugat…
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Jason Ramapuram, Federico Danieli, Eeshan Dhekane +8
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are t…