1 citations · 2 across the 3 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.CL2023
Faithful Model Evaluation for Model-Based Metrics
Palash Goyal, Qian Hu, Rahul Gupta
Statistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they re…
cs.CL2023★ 1 cited
Evaluating Large Language Models on Controlled Generation Tasks
Jiao Sun, Yufei Tian, Wangchunshu Zhou +6
While recent studies have looked into the abilities of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc,…