most citedSuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

4 citations · 4 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2025

COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes

Yunwen Li, Shuangshuang Ying, Xingwei Qu +16

Large language models exhibit systematic deficiencies in creative writing, particularly in non-English contexts where training data is scarce and lacks process-level supervision. W…

cs.CL2025

Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures

Shuangshuang Ying, Yunwen Li, Xingwei Qu +21

Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We intr…

cs.CV2025

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

Yang Chen, Minghao Liu, Yufan Shen +18

The webpage-to-code task requires models to understand visual representations of webpages and generate corresponding code. However, existing benchmarks primarily focus on static sc…

cs.LG2025

Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving

Tianyun Yang, Yunwen Li, Ziniu Li +3

Large vision language models exhibit notable limitations on Geometry Problem Solving (GPS) because of their unreliable diagram interpretation and pure natural-language reasoning. A…

cs.CL20254 cited

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

P Team, Xinrun Du, Yifan Yao +94

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledg…