Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
Ding Chen, Qingchen Yu, Pengyuan Wang +9
With the release of OpenAI's o1 model, reasoning models that adopt slow-thinking strategies have become increasingly common. Their outputs often contain complex reasoning, intermed…
cs.CL2024
Xinyu: An Efficient LLM-based System for Commentary Generation
Yiquan Wu, Bo Tang, Chenyang Xi +13
Commentary provides readers with a deep understanding of events by presenting diverse arguments and evidence. However, creating commentary is a time-consuming task, even for skille…
cs.CL2024
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
Miao Li, Ming-Bin Chen, Bo Tang +8
We present NewsBench, a novel evaluation framework to systematically assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. Our c…