most citedAI Scientists Fail Without Strong Implementation Capability

2 citations · 3 across the 3 of their papers we have counts for

collaborators

9 papers

cs.AI2025

How Far Are AI Scientists from Changing the World?

Qiujie Xie, Yixuan Weng, Minjun Zhu +9

The emergence of large language models (LLMs) is propelling automated scientific discovery to the next level, with LLM-based Artificial Intelligence (AI) Scientist systems now taki…

cs.CL2025

MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs

Shulin Huang, Linyi Yang, Yue Zhang

Large language models exhibit cultural biases and limited cross-cultural understanding capabilities, particularly when serving diverse global user populations. We propose MCEval, a…

cs.AI20252 cited

AI Scientists Fail Without Strong Implementation Capability

Minjun Zhu, Qiujie Xie, Yixuan Weng +4

The emergence of Artificial Intelligence (AI) Scientist represents a paradigm shift in scientific discovery, with large language models (LLMs) taking the lead as the primary execut…

cs.CL2025

DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

Minjun Zhu, Yixuan Weng, Linyi Yang +1

Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face si…

cs.CL2025

ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning

Shulin Huang, Linyi Yang, Yan Song +9

Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challeng…

cs.CL20251 cited

Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values

Hongbo Zhang, Han Cui, Guangsheng Bao +3

We introduce Direct Value Optimization (DVO), an innovative reinforcement learning framework for enhancing large language models in complex reasoning tasks. Unlike traditional meth…