most citedDISGO: Automatic End-to-End Evaluation for Scene Text OCR

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024

Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation

Rastislav Rabatin, Frank Seide, Ernie Chang

We adapt the well-known beam-search algorithm for machine translation to operate in a cascaded real-time speech translation system. This proved to be more complex than initially an…

cs.CL20241 cited

Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time

Frank Seide, Morrie Doulaty, Yangyang Shi +3

We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the f…

eess.AS2024

Effective internal language model training and fusion for factorized transducer model

Jinxi Guo, Niko Moritz, Yingyi Ma +6

The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracte…

eess.AS2024

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

Ju Lin, Niko Moritz, Yiteng Huang +4

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introdu…

cs.CV20232 cited

DISGO: Automatic End-to-End Evaluation for Scene Text OCR

Mei-Yuh Hwang, Yangyang Shi, Ankit Ramchandani +6

This paper discusses the challenges of optical character recognition (OCR) on natural scenes, which is harder than OCR on documents due to the wild content and various image backgr…