cs.CL2023
Interpretable Unified Language Checking
Tianhua Zhang, Hongyin Luo, Yung-Sung Chuang +7
Despite recent concerns about undesirable behaviors generated by large language models (LLMs), including non-factual, biased, and hateful language, we find LLMs are inherent multi-…