5 citations · 12 across the 7 of their papers we have counts for
10 papers · 1 filter
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
Maike Züfle, Ondrej Klejch, Nicholas Sanders +3
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviou…
The Prosody of Emojis
Giulio Zhou, Tsz Kin Lam, Alexandra Birch +1
Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these…
From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM
Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi +5
We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating tex…
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
Tsz Kin Lam, Marco Gaido, Sara Papi +2
Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of com…
Pitfalls and Outlooks in Using COMET
Vilém Zouhar, Pinzhen Chen, Tsz Kin Lam +2
The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a…
Compact Speech Translation Models via Discrete Speech Units Pretraining
Tsz Kin Lam, Alexandra Birch, Barry Haddow
We propose a pretraining method to use Self-Supervised Speech (SSS) model to creating more compact Speech-to-text Translation. In contrast to using the SSS model for initialization…