2 citations · 2 across the 2 of their papers we have counts for
4 papers
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
Kentaro Mitsui, Koh Mitsuda, Toshiaki Wakatsuki +2
Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in resp…
Release of Pre-Trained Models for the Japanese Language
Kei Sawada, Tianyu Zhao, Makoto Shing +5
AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their res…
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
Yukiya Hono, Koh Mitsuda, Tianyu Zhao +3
Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E…
Towards human-like spoken dialogue generation between AI agents from written dialogue
Kentaro Mitsui, Yukiya Hono, Kei Sawada
The advent of large language models (LLMs) has made it possible to generate natural written dialogues between two agents. However, generating human-like spoken dialogues from these…