activity
20192024
most citedToward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

14 citations · 70 across the 21 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models

Zheng Yi Ho, Siyuan Liang, Sen Zhang +2

Hallucinations in Large Language Models (LLMs) remain a major obstacle, particularly in high-stakes applications where factual accuracy is critical. While representation editing an…

cs.CL20233 cited

Unlikelihood Tuning on Negative Samples Amazingly Improves Zero-Shot Translation

Changtong Zan, Liang Ding, Li Shen +4

Zero-shot translation (ZST), which is generally based on a multilingual neural machine translation model, aims to translate between unseen language pairs in training data. The comm…

cs.CL20231 cited

Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State Tracking

Qingyue Wang, Liang Ding, Yanan Cao +5

Zero-shot transfer learning for Dialogue State Tracking (DST) helps to handle a variety of task-oriented dialogue domains without the cost of collecting in-domain data. Existing wo…

cs.CL202214 cited

Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Qihuang Zhong, Liang Ding, Yibing Zhan +11

This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language…

cs.CL2022

TASA: Deceiving Question Answering Models by Twin Answer Sentences Attack

Yu Cao, Dianqi Li, Meng Fang +4

We present Twin Answer Sentences Attack (TASA), an adversarial attack method for question answering (QA) models that produces fluent and grammatical adversarial contexts while main…