6 papers
Multiplayer Interactive World Models with Representation Autoencoders
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents…
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
Chung-Ming Chien, Manu Orsini, Eugene Kharitonov +3
Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time inter…
Continuous Audio Language Models
Simon Rouard, Manu Orsini, Axel Roebel +2
Audio Language Models (ALM) have emerged as the dominant paradigm for speech and music generation by representing audio as sequences of discrete tokens. Yet, unlike text tokens, wh…
Live Music Models
Lyria Team, Antoine Caillon, Brian McWilliams +33
We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release M…
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
Neil Zeghidour, Eugene Kharitonov, Manu Orsini +6
We introduce Delayed Streams Modeling (DSM), a flexible formulation for streaming, multimodal sequence-to-sequence learning. Sequence-to-sequence generation is often cast in an off…
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini +5
We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namel…