2 papers
cs.CL2025
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing
Subramaniam Vincent, Phoebe Wang, Zhan Shi +2
Since the launch of ChatGPT in late 2022, the capacities of Large Language Models and their evaluation have been in constant discussion and evaluation both in academic research and…
cs.CL2024
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
Zhongshen Zeng, Yinhong Liu, Yingjia Wan +16
Large language models (LLMs) have shown increasing capability in problem-solving and decision-making, largely based on the step-by-step chain-of-thought reasoning processes. Howeve…