10 papers
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
Fanqing Meng, Lingxiao Du, Qiguang Chen +4
Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and val…
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Fanqing Meng, Lingxiao Du, Zijian Wu +46
Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator
Alessio Sordo, Lingxiao Du, Meeka-Hanna Lenisa +2
The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, th…
Gym-V: A Unified Vision Environment System for Agentic Vision Research
Fanqing Meng, Lingxiao Du, Jiawei Gu +9
As agentic systems increasingly rely on reinforcement learning from verifiable rewards, standardized ``gym'' infrastructure has become essential for rapid iteration, reproducibilit…
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
Zijian Wu, Xiangyan Liu, Xinyuan Zhang +12
MCP standardizes how LLMs interact with external systems, forming the foundation for general agents. However, existing MCP benchmarks remain narrow in scope: they focus on read-hea…
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
Shanghai AI Lab, :, Yicheng Bao +115
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…