activity
20222024
most citedLIMA: Less Is More for Alignment

128 citations · 251 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20248 cited

OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?

Zhen Huang, Zengzhi Wang, Shijie Xia +1

In this report, we pose the following question: Who is the most intelligent AI model to date, as measured by the OlympicArena (an Olympic-level, multi-discipline, multi-modal bench…

cs.CL20246 cited

Benchmarking Benchmark Leakage in Large Language Models

Ruijie Xu, Zengzhi Wang, Run-Ze Fan +1

Amid the expanding use of pre-training data, the phenomenon of benchmark dataset leakage has become increasingly prominent, exacerbated by opaque training processes and the often u…

cs.CL2024

Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Steffi Chern, Ethan Chern, Graham Neubig +1

Despite the utility of Large Language Models (LLMs) across a wide range of tasks and scenarios, developing a method for reliably evaluating LLMs across varied contexts continues to…

cs.CL202325 cited

FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios

I-Chun Chern, Steffi Chern, Shiqi Chen +6

The emergence of generative pre-trained models has facilitated the synthesis of high-quality text, but it has also posed challenges in identifying factual errors in the generated t…

cs.CL20232 cited

GlobalBench: A Benchmark for Global Progress in Natural Language Processing

Yueqi Song, Catherine Cui, Simran Khanuja +9

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-opt…

cs.CL2023128 cited

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu, Puxin Xu +12

Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and re…