2 papers
cs.CL2025
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
Hai-Chung Nguyen-Phung, Ngoc C. Lê, Van-Chien Nguyen +2
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million death…
cs.CL2025
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
Xinxi Lyu, Yizhong Wang, Hannaneh Hajishirzi +1
Evaluating the capability of Large Language Models (LLMs) in following instructions has heavily relied on a powerful LLM as the judge, introducing unresolved biases that deviate th…