most citedLatent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models

11 citations · 14 across the 5 of their papers we have counts for

collaborators

5 papers

cs.HC20242 cited

No General Code of Ethics for All: Ethical Considerations in Human-bot Psycho-counseling

Lizhi Ma, Tong Zhao, Huachuan Qiu +1

The pervasive use of AI applications is increasingly influencing our everyday decisions. However, the ethical challenges associated with AI transcend conventional ethics and single…

cs.CL20241 cited

Facilitating Pornographic Text Detection for Open-Domain Dialogue Systems via Knowledge Distillation of Large Language Models

Huachuan Qiu, Shuai Zhang, Hongliang He +2

Pornographic content occurring in human-machine interaction dialogues can cause severe side effects for users in open-domain dialogue systems. However, research on detecting pornog…

cs.CL2024

Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents

Shuai Zhang, Yu Lu, Junwen Liu +4

With the growing humanlike nature of dialog agents, people are now engaging in extended conversations that can stretch from brief moments to substantial periods of time. Understand…

cs.CL202311 cited

Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models

Huachuan Qiu, Shuai Zhang, Anqi Li +2

Considerable research efforts have been devoted to ensuring that large language models (LLMs) align with human values and generate safe text. However, an excessive focus on sensiti…

cs.CL2023

A Benchmark for Understanding Dialogue Safety in Mental Health Support

Huachuan Qiu, Tong Zhao, Anqi Li +3

Dialogue safety remains a pervasive challenge in open-domain human-machine interaction. Existing approaches propose distinctive dialogue safety taxonomies and datasets for detectin…