1 paper
Frank Seide, Morrie Doulaty, Yangyang Shi +3
We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the f…