32 citations · 33 across the 13 of their papers we have counts for
4 papers · 1 filter
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation
Sara Papi, Peidong Wang, Junkun Chen +4
The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering transl…
Improving Stability in Simultaneous Speech Translation: A Revision-Controllable Decoding Approach
Junkun Chen, Jian Xue, Peidong Wang +2
Simultaneous Speech-to-Text translation serves a critical role in real-time crosslingual communication. Despite the advancements in recent years, challenges remain in achieving sta…
DiariST: Streaming Speech Translation with Speaker Diarization
Mu Yang, Naoyuki Kanda, Xiaofei Wang +5
End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and han…
Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual Alignments
Sara Papi, Peidong Wang, Junkun Chen +3
In real-world applications, users often require both translations and transcriptions of speech to enhance their comprehension, particularly in streaming scenarios where incremental…