2 citations · 2 across the 3 of their papers we have counts for
3 papers
The Design Space of Tri-Modal Masked Diffusion Models
Louis Bethune, Victor Turrisi, Bruno Kacper Mlodozeniec +21
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal gen…
CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition
Kari A Noriy, Xiaosong Yang, Marcin Budka +1
Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization…
EMNS /Imz/ Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels
Kari Ali Noriy, Xiaosong Yang, Jian Jun Zhang
The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emo…