5 citations · 5 across the 2 of their papers we have counts for
18 papers
Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews
Nuo Chen, Qian Wang, Qingyun Zou +1
When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the diff…
LLM-based Human Simulations Have Not Yet Been Reliable
Qian Wang, Jiaying Wu, Zichen Jiang +6
Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations rema…
CrossAlpha: An Annual-Report Benchmark for Cross-Market Factor Research (with LLM Agents)
Qian Wang, Zhongyi Tong, Nuo Chen +2
Cross-market factor research studies whether firm-level signals from one or more markets can predict returns in a target market, but existing public benchmarks do not support cross…
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Qian Wang, Zhanzhi Lou, Zhenheng Tang +2
LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittl…
LLM DNA: Tracing Model Evolution via Functional Representations
Zhaomin Wu, Haodong Zhao, Ziyang Wang +3
The explosive growth of large language models (LLMs) has created a vast but opaque landscape: millions of models exist, yet their evolutionary relationships through fine-tuning, di…
XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
Nuo Chen, Andre Lin HuiKai, Jiaying Wu +5
Despite the growing adoption of large language models (LLMs) in academic workflows, their capabilities remain limited in supporting high-quality scientific writing. Most existing s…