most citedLLM-based Human Simulations Have Not Yet Been Reliable

5 citations · 5 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

Nuo Chen, Qian Wang, Qingyun Zou +1

When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the diff…

cs.CL20265 cited

LLM-based Human Simulations Have Not Yet Been Reliable

Qian Wang, Jiaying Wu, Zichen Jiang +6

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations rema…

cs.CL2026

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

Qian Wang, Zhanzhi Lou, Zhenheng Tang +2

LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittl…

cs.CL2026

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration

Nuo Chen, Andre Lin HuiKai, Jiaying Wu +5

Despite the growing adoption of large language models (LLMs) in academic workflows, their capabilities remain limited in supporting high-quality scientific writing. Most existing s…

cs.CL2025

JudgeLRM: Large Reasoning Models as a Judge

Nuo Chen, Zhiyuan Hu, Qingyun Zou +4

Large Language Models (LLMs) are increasingly adopted as evaluators, offering a scalable alternative to human annotation. However, existing supervised fine-tuning (SFT) approaches…

cs.CL2025

Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration

Nuo Chen, Yicheng Tong, Jiaying Wu +5

While AI agents show potential in scientific ideation, most existing frameworks rely on single-agent refinement, limiting creativity due to bounded knowledge and perspective. Inspi…