activity
20212023
most citedImproving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives

13 citations · 21 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023★ 6 cited

AutoConv: Automatically Generating Information-seeking Conversations with Large Language Models

Siheng Li, Cheng Yang, Yichun Yin +6

Information-seeking conversation, which aims to help users gather information through conversation, has achieved great progress in recent years. However, the research is still stym…

cs.CL2023★ 2 cited

NewsDialogues: Towards Proactive News Grounded Conversation

Siheng Li, Yichun Yin, Cheng Yang +7

Hot news is one of the most popular topics in daily conversations. However, news grounded conversation has long been stymied by the lack of well-designed task definition and scarce…

cs.AI2022

GUIM -- General User and Item Embedding with Mixture of Representation in E-commerce

Chao Yang, Ru He, Fangquan Lin +3

Our goal is to build general representation (embedding) for each user and each product item across Alibaba's businesses, including Taobao and Tmall which are among the world's bigg…

cs.CL2022★ 13 cited

Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives

Wei Wang, Liangzhu Ge, Jingqiao Zhang +1

Following SimCSE, contrastive learning based methods have achieved the state-of-the-art (SOTA) performance in learning sentence embeddings. However, the unsupervised contrastive le…

cs.CL2021

SAS: Self-Augmentation Strategy for Language Model Pre-training

Yifei Xu, Jingqiao Zhang, Ru He +4

The core of self-supervised learning for pre-training language models includes pre-training task design as well as appropriate data augmentation. Most data augmentations in languag…