activity
20192026
most citedTowards Online End-to-end Transformer Automatic Speech Recognition

31 citations · 53 across the 33 of their papers we have counts for

collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2026

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani +5

Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, t…

cs.CL2026

Benchmarking Speech-to-Speech Translation Models

Alkis Koudounas, Hayato Futami, Quentin Jodelet +3

Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets, preventing direct comparis…

cs.CL2026

Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback

Siddhant Arora, Jinchuan Tian, Jiatong Shi +4

Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single seman…

cs.CL2025

Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems

Siddhant Arora, Jinchuan Tian, Hayato Futami +4

Most end-to-end (E2E) spoken dialogue systems (SDS) rely on voice activity detection (VAD) for turn-taking, but VAD fails to distinguish between pauses and turn completions. Duplex…

cs.CL2025

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

Hayato Futami, Emiru Tsunoo, Yosuke Kashiwagi +4

Speech-to-speech translation (S2ST) has been advanced with large language models (LLMs), which are fine-tuned on discrete speech units. In such approaches, modality adaptation from…

cs.CL2025

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

Yosuke Kashiwagi, Hayato Futami, Emiru Tsunoo +1

This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a di…