9 citations · 42 across the 28 of their papers we have counts for
31 papers · 1 filter
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
Muxi Diao, Yutao Mou, Keqing He +6
The safety of Large Language Models (LLMs) is crucial for the development of trustworthy AI applications. Existing red teaming methods often rely on seed instructions, which limits…
DivTOD: Unleashing the Power of LLMs for Diversifying Task-Oriented Dialogue Representations
Weihao Zeng, Dayuan Fu, Keqing He +3
Language models pre-trained on general text have achieved impressive results in diverse fields. Yet, the distinct linguistic characteristics of task-oriented dialogues (TOD) compar…
Beyond the Known: Investigating LLMs Performance on Out-of-Domain Intent Detection
Pei Wang, Keqing He, Yejie Wang +6
Out-of-domain (OOD) intent detection aims to examine whether the user's query falls outside the predefined domain of the system, which is crucial for the proper functioning of task…
BootTOD: Bootstrap Task-oriented Dialogue Representations by Aligning Diverse Responses
Weihao Zeng, Keqing He, Yejie Wang +2
Pre-trained language models have been successful in many scenarios. However, their usefulness in task-oriented dialogues is limited due to the intrinsic linguistic differences betw…
Multi-Perspective Consistency Enhances Confidence Estimation in Large Language Models
Pei Wang, Yejie Wang, Muxi Diao +3
In the deployment of large language models (LLMs), accurate confidence estimation is critical for assessing the credibility of model predictions. However, existing methods often fa…
Knowledge Editing on Black-box Large Language Models
Xiaoshuai Song, Zhengyang Wang, Keqing He +4
Knowledge editing (KE) aims to efficiently and precisely modify the behavior of large language models (LLMs) to update specific knowledge without negatively influencing other knowl…