papers

Publications (22)

cs.AI2026

RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following

Tianjun Pan, Xuan Lin, Wenyan Yang +7

Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…

cs.CL2026

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

Yin Cai, Zhouhong Gu, Zhaohan Du +5

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly i…

cs.AI2026

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Lujia Zhang, Xingzhou Chen, Hongwei Feng

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory le…

cs.CL2024

StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text

Zhouhong Gu, Haoning Ye, Xingzhou Chen +3

The effective utilization of structured data, integral to corporate data strategies, has been challenged by the rise of large language models (LLMs) capable of processing unstructu…

cs.CL2026

HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns

Xintao Wang, Jian Yang, Weiyuan Li +8

Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona simulation and Role-Playing Langu…

cs.MA2026

Scaling Behavior of Single LLM-Driven Multi-Agent Systems

Jialing Li, Zhouhong Gu, Yin Cai +1

The burgeoning field of LLM-based Multi-Agent Systems (MAS) promises to tackle complex tasks through collaborative intelligence, yet fundamental questions regarding their scaling b…

cs.AI2024

AgentGroupChat: An Interactive Group Chat Simulacra For Better Eliciting Emergent Behavior

Zhouhong Gu, Xiaoxuan Zhu, Haoran Guo +10

Language significantly influences the formation and evolution of Human emergent behavior, which is crucial in understanding collective intelligence within human societies. Consider…

cs.CL2024

Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation

Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye +16

New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive eva…

cs.CL2025

AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need

Zhouhong Gu, Xiaoxuan Zhu, Yin Cai +12

Large language model based multi-agent systems have demonstrated significant potential in social simulation and complex task resolution domains. However, current frameworks face cr…

cs.CL2024

Efficiently Quantifying and Mitigating Ripple Effects in Model Editing

Jianchen Wang, Zhouhong Gu, Xiaoxuan Zhu +5

Large Language Models have revolutionized numerous tasks with their remarkable efficacy. However, editing these models, crucial for rectifying outdated or erroneous information, of…

cs.CL2024

Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models

Zhouhong Gu, Lin Zhang, Jiangjie Chen +8

Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of info…

cs.CL2025

CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs

Jinghao Zhang, Sihang Jiang, Shiwei Guo +7

As large language models (LLMs) are increasingly deployed in diverse cultural environments, evaluating their cultural understanding capability has become essential for ensuring tru…

cs.CL2023

Domain Mastery Benchmark: An Ever-Updating Benchmark for Evaluating Holistic Domain Knowledge of Large Language Model--A Preliminary Release

Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye +7

Domain knowledge refers to the in-depth understanding, expertise, and familiarity with a specific subject, industry, field, or area of special interest. The existing benchmarks are…

cs.CL2025

GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

Zhouhong Gu, Xingzhou Chen, Xiaoran Shi +5

Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to…

cs.MM2025

VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It

Xiaoxuan Zhu, Zhouhong Gu, Sihang Jiang +3

Online courses have significantly lowered the barrier to accessing education, yet the varying content quality of these videos poses challenges. In this work, we focus on the task o…

cs.CL2024

DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?

Zhouhong Gu, Lin Zhang, Xiaoxuan Zhu +8

Detecting evidence within the context is a key step in the process of reasoning task. Evaluating and enhancing the capabilities of LLMs in evidence detection will strengthen contex…

cs.CL2023

Sem4SAP: Synonymous Expression Mining From Open Knowledge Graph For Language Model Synonym-Aware Pretraining

Zhouhong Gu, Sihang Jiang, Wenhao Huang +3

The model's ability to understand synonymous expression is crucial in many kinds of downstream tasks. It will make the model to better understand the similarity between context, an…

cs.CL2025

RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model

Lin Zhang, Zhouhong Gu, Xiaoran Shi +2

As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitation…

cs.CL2024

LLM-GAN: Construct Generative Adversarial Network Through Large Language Models For Explainable Fake News Detection

Yifeng Wang, Zhouhong Gu, Siwei Zhang +5

Explainable fake news detection predicts the authenticity of news items with annotated explanations. Today, Large Language Models (LLMs) are known for their powerful natural langua…

cs.AI2023

GANTEE: Generative Adversatial Network for Taxonomy Entering Evaluation

Zhouhong Gu, Sihang Jiang, Jingping Liu +5

Taxonomy is formulated as directed acyclic concepts graphs or trees that support many downstream tasks. Many new coming concepts need to be added to an existing taxonomy. The tradi…

cs.CL2025

LITE: LLM-Impelled efficient Taxonomy Evaluation

Lin Zhang, Zhouhong Gu, Suhang Zheng +4

This paper presents LITE, an LLM-based evaluation method designed for efficient and flexible assessment of taxonomy quality. To address challenges in large-scale taxonomy evaluatio…

cs.CL2025

ToReMi: Topic-Aware Data Reweighting for Dynamic Pre-Training Data Selection

Xiaoxuan Zhu, Zhouhong Gu, Baiqian Wu +5

Pre-training large language models (LLMs) necessitates enormous diverse textual corpora, making effective data selection a key challenge for balancing computational resources and m…