14 papers
Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews
Nuo Chen, Qian Wang, Qingyun Zou +1
When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the diff…
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Qian Wang, Zhanzhi Lou, Zhenheng Tang +2
LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittl…
Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation
Nuo Chen, Yicheng Tong, Yuzhe Yang +5
Multi-agent systems (MAS) are increasingly used for open-ended idea generation, driven by the expectation that collective interaction will broaden the exploration diversity. Howeve…
MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
Qingyun Zou, Jiahao Cui, Nuo Chen +2
Large language models (LLMs) have achieved strong performance on code completion tasks in general-purpose programming languages. However, existing repository-level code completion…
PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing
Junyi Hou, Andre Lin Huikai, Nuo Chen +2
Large language models are increasingly embedded into academic writing workflows, yet existing assistants remain external to the editor, preventing deep interaction with document st…
HLStrans: Dataset for C-to-HLS Hardware Code Synthesis
Qingyun Zou, Nuo Chen, Yao Chen +2
High-Level Synthesis (HLS) enables hardware design from C/C++ kernels but requires extensive transformations, such as restructuring code, inserting pragmas, adapting data types, an…