Publications (22)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
Tianjun Pan, Xuan Lin, Wenyan Yang +7
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…
MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments
Yin Cai, Zhouhong Gu, Zhaohan Du +5
Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly i…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
Lujia Zhang, Xingzhou Chen, Hongwei Feng
Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory le…
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
Zhouhong Gu, Haoning Ye, Xingzhou Chen +3
The effective utilization of structured data, integral to corporate data strategies, has been challenged by the rise of large language models (LLMs) capable of processing unstructu…
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
Xintao Wang, Jian Yang, Weiyuan Li +8
Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona simulation and Role-Playing Langu…
Scaling Behavior of Single LLM-Driven Multi-Agent Systems
Jialing Li, Zhouhong Gu, Yin Cai +1
The burgeoning field of LLM-based Multi-Agent Systems (MAS) promises to tackle complex tasks through collaborative intelligence, yet fundamental questions regarding their scaling b…
AgentGroupChat: An Interactive Group Chat Simulacra For Better Eliciting Emergent Behavior
Zhouhong Gu, Xiaoxuan Zhu, Haoran Guo +10
Language significantly influences the formation and evolution of Human emergent behavior, which is crucial in understanding collective intelligence within human societies. Consider…
Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye +16
New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive eva…
AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need
Zhouhong Gu, Xiaoxuan Zhu, Yin Cai +12
Large language model based multi-agent systems have demonstrated significant potential in social simulation and complex task resolution domains. However, current frameworks face cr…
Efficiently Quantifying and Mitigating Ripple Effects in Model Editing
Jianchen Wang, Zhouhong Gu, Xiaoxuan Zhu +5
Large Language Models have revolutionized numerous tasks with their remarkable efficacy. However, editing these models, crucial for rectifying outdated or erroneous information, of…
Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models
Zhouhong Gu, Lin Zhang, Jiangjie Chen +8
Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of info…
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
Jinghao Zhang, Sihang Jiang, Shiwei Guo +7
As large language models (LLMs) are increasingly deployed in diverse cultural environments, evaluating their cultural understanding capability has become essential for ensuring tru…
Domain Mastery Benchmark: An Ever-Updating Benchmark for Evaluating Holistic Domain Knowledge of Large Language Model--A Preliminary Release
Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye +7
Domain knowledge refers to the in-depth understanding, expertise, and familiarity with a specific subject, industry, field, or area of special interest. The existing benchmarks are…
GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization
Zhouhong Gu, Xingzhou Chen, Xiaoran Shi +5
Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to…
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
Xiaoxuan Zhu, Zhouhong Gu, Sihang Jiang +3
Online courses have significantly lowered the barrier to accessing education, yet the varying content quality of these videos poses challenges. In this work, we focus on the task o…
DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?
Zhouhong Gu, Lin Zhang, Xiaoxuan Zhu +8
Detecting evidence within the context is a key step in the process of reasoning task. Evaluating and enhancing the capabilities of LLMs in evidence detection will strengthen contex…
Sem4SAP: Synonymous Expression Mining From Open Knowledge Graph For Language Model Synonym-Aware Pretraining
Zhouhong Gu, Sihang Jiang, Wenhao Huang +3
The model's ability to understand synonymous expression is crucial in many kinds of downstream tasks. It will make the model to better understand the similarity between context, an…
RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model
Lin Zhang, Zhouhong Gu, Xiaoran Shi +2
As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitation…
LLM-GAN: Construct Generative Adversarial Network Through Large Language Models For Explainable Fake News Detection
Yifeng Wang, Zhouhong Gu, Siwei Zhang +5
Explainable fake news detection predicts the authenticity of news items with annotated explanations. Today, Large Language Models (LLMs) are known for their powerful natural langua…
GANTEE: Generative Adversatial Network for Taxonomy Entering Evaluation
Zhouhong Gu, Sihang Jiang, Jingping Liu +5
Taxonomy is formulated as directed acyclic concepts graphs or trees that support many downstream tasks. Many new coming concepts need to be added to an existing taxonomy. The tradi…
LITE: LLM-Impelled efficient Taxonomy Evaluation
Lin Zhang, Zhouhong Gu, Suhang Zheng +4
This paper presents LITE, an LLM-based evaluation method designed for efficient and flexible assessment of taxonomy quality. To address challenges in large-scale taxonomy evaluatio…
ToReMi: Topic-Aware Data Reweighting for Dynamic Pre-Training Data Selection
Xiaoxuan Zhu, Zhouhong Gu, Baiqian Wu +5
Pre-training large language models (LLMs) necessitates enormous diverse textual corpora, making effective data selection a key challenge for balancing computational resources and m…