4 papers
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
Yongan Yu, Qingchen Hu, Xianda Du +3
Climate change adaptation requires the understanding of disruptive weather impacts on society, where large language models (LLMs) might be applicable. However, their effectiveness…
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
Enze Zhang, Jiaying Wang, Mengxi Xiao +7
Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on su…
Long Term Memory: The Foundation of AI Self-Evolution
Xun Jiang, Feng Li, Han Zhao +12
Large language models (LLMs) like GPTs, trained on vast datasets, have demonstrated impressive capabilities in language understanding, reasoning, and planning, achieving human-leve…
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
Jianfei He, Lilin Wang, Jiaying Wang +5
Identifying offensive language is essential for maintaining safety and sustainability in the social media era. Though large language models (LLMs) have demonstrated encouraging pot…