3 citations · 10 across the 29 of their papers we have counts for
Showing 2023 · cs.CLShow all
2 papers · 2 filters
cs.CL2023★ 2 cited
Flames: Benchmarking Value Alignment of LLMs in Chinese
Kexin Huang, Xiangyang Liu, Qianyu Guo +9
The widespread adoption of large language models (LLMs) across various regions underscores the urgent need to evaluate their alignment with human values. Current benchmarks, howeve…
cs.CL2023★ 1 cited
Fake Alignment: Are LLMs Really Aligned Well?
Yixu Wang, Yan Teng, Kexin Huang +7
The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored is…