6 papers · 1 filter
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
Zenan Zhai, Hao Li, Xudong Han +4
Recent advances in large language models (LLMs) have shown that they can answer questions requiring complex reasoning. However, their ability to identify and respond to text contai…
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
Renxi Wang, Honglin Mu, Liqun Ma +5
Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchma…
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
Maiya Goloburda, Nurkhan Laiyk, Diana Turmakhan +11
Large language models (LLMs) are known to have the potential to generate harmful content, posing risks to users. While significant progress has been made in developing taxonomies f…
Loki: An Open-Source Tool for Fact Verification
Haonan Li, Xudong Han, Hao Wang +7
We introduce Loki, an open-source tool designed to address the growing problem of misinformation. Loki adopts a human-centered approach, striking a balance between the quality of f…
Arabic Dataset for LLM Safeguard Evaluation
Yasser Ashraf, Yuxia Wang, Bin Gu +2
The growing use of large language models (LLMs) has raised concerns regarding their safety. While many studies have focused on English, the safety of LLMs in Arabic, with its lingu…
ToolGen: Unified Tool Retrieval and Calling via Generation
Renxi Wang, Xudong Han, Lei Ji +3
As large language models (LLMs) advance, their inability to autonomously execute tasks by directly interacting with external tools remains a critical limitation. Traditional method…