3 papers
cs.LG2024
Backtracking Improves Generation Safety
Yiming Zhang, Jianfeng Chi, Hailey Nguyen +4
Text generation has a fundamental limitation almost by definition: there is no taking back tokens that have been generated, even when they are clearly problematic. In the context o…
cs.CL2024
Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models
Yi-Lin Tuan, Xilun Chen, Eric Michael Smith +5
As large language models (LLMs) become easily accessible nowadays, the trade-off between safety and helpfulness can significantly impact user experience. A model that prioritizes s…
cs.CL2023
Step by Step to Fairness: Attributing Societal Bias in Task-oriented Dialogue Systems
Hsuan Su, Rebecca Qian, Chinnadhurai Sankar +4
Recent works have shown considerable improvements in task-oriented dialogue (TOD) systems by utilizing pretrained large language models (LLMs) in an end-to-end manner. However, the…