2 papers
cs.CV2023
Self-Enhancement Improves Text-Image Retrieval in Foundation Visual-Language Models
Yuguang Yang, Yiming Wang, Shupeng Geng +4
The emergence of cross-modal foundation models has introduced numerous approaches grounded in text-image retrieval. However, on some domain-specific retrieval tasks, these models f…
eess.AS2023
HYBRIDFORMER: improving SqueezeFormer with hybrid attention and NSR mechanism
Yuguang Yang, Yu Pan, Jingjing Yin +3
SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attenti…