3 papers
cs.CV2024★ 4 cited
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
Xingning Dong, Qingpei Guo, Tian Gan +5
We present a framework for learning cross-modal video representations by directly pre-training on raw data to facilitate various downstream video-text tasks. Our main contributions…
cs.CL2023★ 32 cited
A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition
Ruchao Fan, Wei Chu, Peng Chang +1
Recently, end-to-end models have been widely used in automatic speech recognition (ASR) systems. Two of the most representative approaches are connectionist temporal classification…
cs.CV2023★ 3 cited
DC-Former: Diverse and Compact Transformer for Person Re-Identification
Wen Li, Cheng Zou, Meng Wang +5
In person re-identification (re-ID) task, it is still challenging to learn discriminative representation by deep learning, due to limited data. Generally speaking, the model will g…