5 papers
BABE: Biology Arena BEnchmark
Junting Zhou, Jin Chen, Linfeng Hao +10
The rapid evolution of large language models (LLMs) has expanded their capabilities from basic dialogue to advanced scientific reasoning. However, existing benchmarks in biology of…
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
Yiming Liang, Yizhi Li, Yantao Du +14
Benchmarks play a crucial role in tracking the rapid advancement of large language models (LLMs) and identifying their capability boundaries. However, existing benchmarks predomina…
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong +31
Existing Vision-Language Models often struggle with complex, multi-question reasoning tasks where partial correctness is crucial for effective learning. Traditional reward mechanis…
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
Luoxin Chen, Jinming Gu, Liankai Huang +33
LLMs have demonstrated strong mathematical reasoning abilities by leveraging reinforcement learning with long chain-of-thought, yet they continue to struggle with theorem proving d…
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
Zhongyuan Peng, Yifan Yao, Kaijing Ma +16
Translating natural language mathematical statements into formal, executable code is a fundamental challenge in automated theorem proving. While prior work has focused on generatio…