papers

Publications (32)

cs.CL2022

Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt Tuning

Ziyun Xu, Chengyu Wang, Minghui Qiu +4

Pre-trained Language Models (PLMs) have achieved remarkable performance for various language understanding tasks in IR systems, which require the fine-tuning process based on label…

cs.CL2021

Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning

Runxin Xu, Fuli Luo, Zhiyuan Zhang +4

Recent pretrained language models extend from millions to billions of parameters. Thus the need to fine-tune an extremely large pretrained model with a limited training corpus aris…

cs.CL2024

DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Huajian Xin, Z. Z. Ren, Junxiao Song +14

We introduce DeepSeek-Prover-V1.5, an open-source language model designed for theorem proving in Lean 4, which enhances DeepSeek-Prover-V1 by optimizing both training and inference…

cs.CL2020

CAPT: Contrastive Pre-Training for Learning Denoised Sequence Representations

Fuli Luo, Pengcheng Yang, Shicheng Li +2

Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing. These models typica…

cs.CL2021

VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation

Fuli Luo, Wei Wang, Jiahao Liu +5

Existing work in multilingual pretraining has demonstrated the potential of cross-lingual transferability by training a unified Transformer encoder for multiple languages. However,…

cs.CL2021

SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels

Chenliang Li, Ming Yan, Haiyang Xu +4

Vision-language pre-training (VLP) on large-scale image-text pairs has recently witnessed rapid progress for learning cross-modal representations. Existing pre-training methods eit…