1 citations · 1 across the 4 of their papers we have counts for
6 papers
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
Zerui Cheng, Jiashuo Liu, Jianzhu Yao +3
Standard tabular benchmarks mainly focus on the evaluation of a model's capability to interpolate values inside a data manifold, where models good at performing local statistical s…
MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
Kaiyuan Zhang, Chenghao Yang, Zhoufutu Wen +19
As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behav…
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
Zeng Zhiyuan, Jiashuo Liu, Zhangyue Yin +3
While Reinforcement Learning for Verifiable Rewards (RLVR) is powerful for training large reasoning models, its training dynamics harbor a critical challenge: RL overfitting, where…
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
Ziniu Li, Congliang Chen, Tianyun Yang +5
Large Language Models (LLMs) can self-improve through reinforcement learning, where they generate trajectories to explore and discover better solutions. However, this exploration p…
SciDA: Scientific Dynamic Assessor of LLMs
Junting Zhou, Tingjia Miao, Yiyan Liao +15
Advancement in Large Language Models (LLMs) reasoning capabilities enables them to solve scientific problems with enhanced efficacy. Thereby, a high-quality benchmark for comprehen…
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
Chenghao Yang, Yinbo Luo, Zhoufutu Wen +8
Large Language Models (\textbf{LLMs}), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs' robustness, especially in handling long complex di…