3 papers
cs.CL2026
BenchBench: Benchmarking Automated Benchmark Generation
Yandan Zheng, Haoran Luo, Zhenghong Lin +2
Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are c…
cs.AI2026
OrchMAS: Orchestrated Reasoning with Multi Collaborative Heterogeneous Scientific Expert Structured Agents
Yichao Feng, Haoran Luo, Zhenghong Lin +4
Multi-agent large language model frameworks are promising for complex multi step reasoning, yet existing systems remain weak for scientific and knowledge intensive domains due to s…
cs.CL2026
SSL: Sweet Spot Learning for Differentiated Guidance in Agentic Optimization
Jinyang Wu, Changpeng Yang, Yuhao Shen +9
Reinforcement learning with verifiable rewards has emerged as a powerful paradigm for training intelligent agents. However, existing methods typically employ binary rewards that fa…