Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
arXiv:2107.03451
Abstract
Over the last several years, end-to-end neural conversational agents have vastly improved in their ability to carry a chit-chat conversation with humans. However, these models are often trained on large datasets from the internet, and as a result, may learn undesirable behaviors from this data, such as toxic or otherwise harmful language. Researchers must thus wrestle with the issue of how and when to release these models. In this paper, we survey the problem landscape for safety for end-to-end conversational AI and discuss recent and related work. We highlight tensions between values, potential positive impact and potential harms, and provide a framework for making decisions about whether and how to release these models, following the tenets of value-sensitive design. We additionally provide a suite of tools to enable researchers to make better-informed decisions about training and releasing end-to-end conversational AI models.
References in corpus (13)
- Language Models are Few-Shot Learners
- Towards a Human-like Open-Domain Chatbot
- Conversational AI: The Science Behind the Alexa Prize
- Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge
- Plug and Play Language Models: A Simple Approach to Controlled Text Generation
- Advancing the State of the Art in Open Domain Dialog Systems through the Alexa Prize
- SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval)
- Neural Generation Meets Real People: Towards Emotionally Engaging Mixed-Initiative Conversations
- Chat as Expected: Learning to Manipulate Black-box Neural Dialogue Models
- Challenges in Automated Debiasing for Toxic Language Detection
- Detoxifying Language Models Risks Marginalizing Minority Voices
- Deploying Lifelong Open-Domain Dialogue Learning
- Detecting and Classifying Malevolent Dialogue Responses: Taxonomy, Data and Methodology