activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

OpenCompass: A Universal Evaluation Platform for Large Language Models

Maosong Cao, Kai Chen, Haodong Duan +27

In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…

cs.CL2025

NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities

Mo Li, Songyang Zhang, Taolin Zhang +3

The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-…

cs.CL2025

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

Taolin Zhang, Maosong Cao, Alexander Lam +2

Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness…

cs.CL2025

Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement

Maosong Cao, Taolin Zhang, Mo Li +5

The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs). However, as LLMs become more ad…

cs.CL2024

CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Maosong Cao, Alexander Lam, Haodong Duan +3

Efficient and accurate evaluation is crucial for the continuous improvement of large language models (LLMs). Among various assessment methods, subjective evaluation has garnered si…

cs.CL2024

HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

Haoran Que, Feiyu Duan, Liqun He +11

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks (e.g., long-context understanding), and many benchmarks have been proposed.…