1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 1 cited
InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu +5
The local and global features are both essential for automatic speech recognition (ASR). Many recent methods have verified that simply combining local and global features can furth…
cs.CL2023
Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Hai-Bo Qin, Zhi-Hao Lai +5
Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and seman…
cs.CV2023
MRET: Multi-resolution Transformer for Video Quality Assessment
Junjie Ke, Tianhao Zhang, Yilin Wang +2
No-reference video quality assessment (NR-VQA) for user generated content (UGC) is crucial for understanding and improving visual experience. Unlike video recognition tasks, VQA ta…