Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
Chuang Liu, Linhao Yu, Jiaxuan Li +11
The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluat…
cs.CL2023
RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models
Tianhao Shen, Sun Li, Quan Tu +1
The rapid evolution of large language models necessitates effective benchmarks for evaluating their role knowledge, which is essential for establishing connections with the real wo…