12 citations · 20 across the 10 of their papers we have counts for
6 papers
InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu +5
The local and global features are both essential for automatic speech recognition (ASR). Many recent methods have verified that simply combining local and global features can furth…
Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Hai-Bo Qin, Zhi-Hao Lai +5
Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and seman…
VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision
Mengyin Liu, Jie Jiang, Chao Zhu +1
Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects ofte…
Arbitrary Shape Text Detection via Segmentation with Probability Maps
Shi-Xue Zhang, Xiaobin Zhu, Lei Chen +2
Arbitrary shape text detection is a challenging task due to the significantly varied sizes and aspect ratios, arbitrary orientations or shapes, inaccurate annotations, etc. Due to…
Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification
Zhiyu Fang, Xiaobin Zhu, Chun Yang +3
Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due…
Learning to Diversify via Weighted Kernels for Classifier Ensemble
Xu-Cheng Yin, Chun Yang, Hong-Wei Hao
Classifier ensemble generally should combine diverse component classifiers. However, it is difficult to give a definitive connection between diversity measure and ensemble accuracy…