4 papers
ParlAI Vote: A Web Platform for Analyzing Gender and Political Bias in Large Language Models
Wenjie Lin, Hange Liu, Yingying Zhuang +5
We present ParlAI Vote, an interactive web platform for exploring European Parliament debates and votes, and for testing LLMs on vote prediction and bias analysis. This web system…
Benchmarking Gender and Political Bias in Large Language Models
Jinrui Yang, Xudong Han, Timothy Baldwin
We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-cal…
ToolGen: Unified Tool Retrieval and Calling via Generation
Renxi Wang, Xudong Han, Lei Ji +3
As large language models (LLMs) advance, their inability to autonomously execute tasks by directly interacting with external tools remains a critical limitation. Traditional method…
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
Zenan Zhai, Hao Li, Xudong Han +4
Recent advances in large language models (LLMs) have shown that they can answer questions requiring complex reasoning. However, their ability to identify and respond to text contai…