1 citations · 2 across the 2 of their papers we have counts for
5 papers
How Catastrophic is Your LLM? Certifying Risk in Conversation
Chengxiao Wang, Isha Chaudhary, Qian Hu +3
Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to…
Certifying Counterfactual Bias in LLMs
Isha Chaudhary, Qian Hu, Manoj Kumar +3
Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across…
Toward Informal Language Processing: Knowledge of Slang in Large Language Models
Zhewei Sun, Qian Hu, Rahul Gupta +2
Recent advancement in large language models (LLMs) has offered a strong potential for natural language systems to process informal language. A representative form of informal langu…
Faithful Model Evaluation for Model-Based Metrics
Palash Goyal, Qian Hu, Rahul Gupta
Statistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they re…
Evaluating Large Language Models on Controlled Generation Tasks
Jiao Sun, Yufei Tian, Wangchunshu Zhou +6
While recent studies have looked into the abilities of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc,…