1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 1 cited
i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data
Ziyi Yang, Mahmoud Khademi, Yichong Xu +16
The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…
cs.CL2021
Sequence-level self-learning with multiple hypotheses
Kenichi Kumatani, Dimitrios Dimitriadis, Yashesh Gaur +4
In this work, we develop new self-learning techniques with an attention-based sequence-to-sequence (seq2seq) model for automatic speech recognition (ASR). For untranscribed speech…