2 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 2 cited
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
Liang Xu, Hang Xue, Lei Zhu +1
We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chi…
cs.CL2023★ 2 cited
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese
Liang Xu, Kangkang Zhao, Lei Zhu +1
Large language models (LLMs), like ChatGPT and GPT-4, have demonstrated remarkable abilities in natural language understanding and generation. However, alongside their positive imp…