31 citations · 31 across the 2 of their papers we have counts for
2 papers
cs.CL2023
Massively Multilingual Shallow Fusion with Large Language Models
Ke Hu, Tara N. Sainath, Bo Li +7
While large language models (LLM) have made impressive progress in natural language processing, it remains unclear how to utilize them in improving automatic speech recognition (AS…
cs.CV2021★ 31 cited
Co-training Transformer with Videos and Images Improves Action Recognition
Bowen Zhang, Jiahui Yu, Christopher Fifty +4
In learning action recognition, models are typically pre-trained on object recognition with images, such as ImageNet, and later fine-tuned on target action recognition with videos.…