7 papers · 1 filter
OpenCompass: A Universal Evaluation Platform for Large Language Models
Maosong Cao, Kai Chen, Haodong Duan +27
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
Mo Li, Songyang Zhang, Taolin Zhang +3
The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-…
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
Taolin Zhang, Maosong Cao, Alexander Lam +2
Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness…
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
Maosong Cao, Taolin Zhang, Mo Li +5
The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs). However, as LLMs become more ad…
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
Maosong Cao, Alexander Lam, Haodong Duan +3
Efficient and accurate evaluation is crucial for the continuous improvement of large language models (LLMs). Among various assessment methods, subjective evaluation has garnered si…
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
Haoran Que, Feiyu Duan, Liqun He +11
In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks (e.g., long-context understanding), and many benchmarks have been proposed.…