activity
20172025
most citedWhy We Need New Evaluation Metrics for NLG

157 citations · 158 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2025★ 1 cited

ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models

Boyang Xue, Qi Zhu, Rui Wang +8

Although demonstrating remarkable performance on reasoning tasks, Large Language Models (LLMs) still tend to fabricate unreliable responses when confronted with problems that are u…

cs.CL2025

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Rui Wang, Hongru Wang, Boyang Xue +9

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to perform complex reasoning tasks, transitioning from fast and intuitive thinking (Sy…

cs.CL2024

MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models

Boyang Xue, Hongru Wang, Rui Wang +5

The tendency of Large Language Models (LLMs) to generate hallucinations raises concerns regarding their reliability. Therefore, confidence estimations indicating the extent of trus…

cs.CL2020

CUHK at SemEval-2020 Task 4: CommonSense Explanation, Reasoning and Prediction with Multi-task Learning

Hongru Wang, Xiangru Tang, Sunny Lai +4

This paper describes our system submitted to task 4 of SemEval 2020: Commonsense Validation and Explanation (ComVE) which consists of three sub-tasks. The task is to directly valid…

cs.CL2017★ 157 cited

Why We Need New Evaluation Metrics for NLG

Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry +1

The majority of NLG evaluation relies on automatic metrics, such as BLEU . In this paper, we motivate the need for novel, system- and data-independent automatic evaluation methods:…