1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.SD2024
Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
Bin Wu, Shinnosuke Takamichi, Sakriani Sakti +1
The marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, a…
cs.CL2024★ 1 cited
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
Xi Chen, Songyang Zhang, Qibing Bai +2
We introduces LLaST, a framework for building high-performance Large Language model based Speech-to-text Translation systems. We address the limitations of end-to-end speech transl…
cs.CL2024
NAIST Simultaneous Speech Translation System for IWSLT 2024
Yuka Ko, Ryo Fukuda, Yuta Nishikawa +9
This paper describes NAIST's submission to the simultaneous track of the IWSLT 2024 Evaluation Campaign: English-to-{German, Japanese, Chinese} speech-to-text translation and Engli…