10 papers
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
Zejun Liu, Jian Wu, Ru Peng +4
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focus…
League: Leaderboard Generation on Demand
Jian Wu, Jiayu Zhang, Dongyuan Li +5
This paper introduces Leaderboard Auto Generation (LAG), a novel and well-organized framework for automatic generation of leaderboards on a given research topic in rapidly evolving…
How Far Are AI Scientists from Changing the World?
Qiujie Xie, Yixuan Weng, Minjun Zhu +9
The emergence of large language models (LLMs) is propelling automated scientific discovery to the next level, with LLM-based Artificial Intelligence (AI) Scientist systems now taki…
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
Shulin Huang, Linyi Yang, Yue Zhang
Large language models exhibit cultural biases and limited cross-cultural understanding capabilities, particularly when serving diverse global user populations. We propose MCEval, a…
Constrain Alignment with Sparse Autoencoders
Qingyu Yin, Chak Tou Leong, Minjun Zhu +7
The alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF)…
AI Scientists Fail Without Strong Implementation Capability
Minjun Zhu, Qiujie Xie, Yixuan Weng +4
The emergence of Artificial Intelligence (AI) Scientist represents a paradigm shift in scientific discovery, with large language models (LLMs) taking the lead as the primary execut…