3 papers
cs.LG2026
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4
We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but…
cs.LG2026
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Artem Ploujnikov, Francesco Verdini, Samir Sadok +1
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). How…
cs.CL2025
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
Francesco Verdini, Pierfrancesco Melucci, Stefano Perna +9
The remarkable performance achieved by Large Language Models (LLM) has driven research efforts to leverage them for a wide range of tasks and input modalities. In speech-to-text (S…