765 citations · 852 across the 13 of their papers we have counts for
15 papers
Back-Translation-Style Data Augmentation for Mandarin Chinese Polyphone Disambiguation
Chunyu Qiang, Peng Yang, Hao Che +3
Conversion of Chinese Grapheme-to-Phoneme (G2P) plays an important role in Mandarin Chinese Text-To-Speech (TTS) systems, where one of the biggest challenges is the task of polypho…
InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings
Xing Wu, Chaochen Gao, Zijia Lin +3
Contrastive learning has been extensively studied in sentence embedding learning, which assumes that the embeddings of different views of the same sentence are closer. The constrai…
RaP: Redundancy-aware Video-language Pre-training for Text-Video Retrieval
Xing Wu, Chaochen Gao, Zijia Lin +3
Video language pre-training methods have mainly adopted sparse sampling techniques to alleviate the temporal redundancy of videos. Though effective, sparse sampling still suffers i…
TokenFlow: Rethinking Fine-grained Cross-modal Alignment in Vision-Language Retrieval
Xiaohan Zou, Changqiao Wu, Lele Cheng +1
Most existing methods in vision-language retrieval match two modalities by either comparing their global feature vectors which misses sufficient information and lacks interpretabil…
Text Smoothing: Enhance Various Data Augmentation Methods on Text Classification Tasks
Xing Wu, Chaochen Gao, Meng Lin +3
Before entering the neural network, a token is generally converted to the corresponding one-hot representation, which is a discrete distribution of the vocabulary. Smoothed represe…
MlTr: Multi-label Classification with Transformer
Xing Cheng, Hezheng Lin, Xiangyu Wu +5
The task of multi-label image classification is to recognize all the object labels presented in an image. Though advancing for years, small objects, similar objects and objects wit…