papers

Publications (12)

cs.CL2025

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Zhen Huang, Zengzhi Wang, Shijie Xia +25

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showc…

cs.CL2024

Reformatted Alignment

Run-Ze Fan, Xuefeng Li, Haoyang Zou +5

The quality of finetuning data is crucial for aligning large language models (LLMs) with human values. Current methods to improve data quality are either labor-intensive or prone t…

cs.CL2025

MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning

Run-Ze Fan, Zengzhi Wang, Pengfei Liu

Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source com…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.CL2024

Benchmarking Benchmark Leakage in Large Language Models

Ruijie Xu, Zengzhi Wang, Run-Ze Fan +1

Amid the expanding use of pre-training data, the phenomenon of benchmark dataset leakage has become increasingly prominent, exacerbated by opaque training processes and the often u…

cs.CL2025

Deep Research: A Systematic Survey

Zhengliang Shi, Yiqun Chen, Haitao Li +23

Large language models (LLMs) have rapidly evolved from text generators into powerful problem solvers. Yet, many open tasks demand critical thinking, multi-source, and verifiable ou…

cs.CL2024

Data Contamination Report from the 2024 CONDA Shared Task

Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25

The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…

cs.CL2023

MerA: Merging Pretrained Adapters For Few-Shot Learning

Shwai He, Run-Ze Fan, Liang Ding +3

Adapter tuning, which updates only a few parameters, has become a mainstream method for fine-tuning pretrained language models to downstream tasks. However, it often yields subpar…

cs.CL2023

Merging Experts into One: Improving Computational Efficiency of Mixture of Experts

Shwai He, Run-Ze Fan, Liang Ding +3

Scaling the size of language models usually leads to remarkable advancements in NLP tasks. But it often comes with a price of growing computational cost. Although a sparse Mixture…

cs.CL2025

Generative AI Act II: Test Time Scaling Drives Cognition Engineering

Shijie Xia, Yiwei Qin, Xuefeng Li +11

The first generation of Large Language Models - what might be called "Act I" of generative AI (2020-2023) - achieved remarkable success through massive parameter and data scaling,…

cs.CL2023

RIGHT: Retrieval-augmented Generation for Mainstream Hashtag Recommendation

Run-Ze Fan, Yixing Fan, Jiangui Chen +3

Automatic mainstream hashtag recommendation aims to accurately provide users with concise and popular topical hashtags before publication. Generally, mainstream hashtag recommendat…

cs.CL2023

Generative Judge for Evaluating Alignment

Junlong Li, Shichao Sun, Weizhe Yuan +3

The rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers…