4 papers
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Liya Zhu, Xin Ma, Tao Liu +35
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on re…
Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios
Tao Liu, Ye Lu, Ruohua Zhang +4
Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depe…
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
Xin Lai, Junyi Li, Wei Li +3
Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing open-source approaches often ex…
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Jingkai He, Tianjian Li, Erhu Feng +5
With the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlik…