activity
20242026
collaborators

5 papers

cs.SD2026

MuScriptor: An Open Model for Multi-Instrument Music Transcription

Simon Rouard, Michael Krause, Axel Roebel +2

Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic…

cs.CL2026

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez +1

Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely wi…

cs.CV2025

Vision-Speech Models: Teaching Speech Models to Converse about Images

Amélie Royer, Moritz Böhle, Gabriel de Marmiesse +4

The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards b…

cs.CL2025

High-Fidelity Simultaneous Speech-To-Speech Translation

Tom Labiausse, Laurent Mazaré, Edouard Grave +3

We introduce Hibiki, a decoder-only model for simultaneous speech translation. Hibiki leverages a multistream language model to synchronously process source and target speech, and…

eess.AS2024

Moshi: a speech-text foundation model for real-time dialogue

Alexandre Défossez, Laurent Mazaré, Manu Orsini +5

We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namel…