most citedAssessing Judging Bias in Large Reasoning Models: An Empirical Study

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CL2025

Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration

Nuo Chen, Yicheng Tong, Jiaying Wu +5

While AI agents show potential in scientific ideation, most existing frameworks rely on single-agent refinement, limiting creativity due to bounded knowledge and perspective. Inspi…

cs.CY2025

Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference

Nuo Chen, Moming Duan, Andre Huikai Lin +3

Artificial Intelligence (AI) conferences are essential for advancing research, sharing knowledge, and fostering academic community. However, their rapid expansion has rendered the…

cs.CY2025

Towards Evaluting Fake Reasoning Bias in Language Models

Qian Wang, Zhenheng Tang, Zhanzhi Lou +3

Large Reasoning Models (LRMs), evolved from standard Large Language Models (LLMs), are increasingly utilized as automated judges because of their explicit reasoning processes. Yet…

cs.CY20251 cited

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Qian Wang, Zhanzhi Lou, Zhenheng Tang +5

Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge s…

cs.CL2025

JudgeLRM: Large Reasoning Models as a Judge

Nuo Chen, Zhiyuan Hu, Qingyun Zou +4

Large Language Models (LLMs) are increasingly adopted as evaluators, offering a scalable alternative to human annotation. However, existing supervised fine-tuning (SFT) approaches…

cs.CY2025

From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?

Qian Wang, Zhenheng Tang, Bingsheng He

Simulation powered by Large Language Models (LLMs) has become a promising method for exploring complex human social behaviors. However, the application of LLMs in simulations prese…