activity
20192025
most citedJais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

23 citations · 64 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL2025

Benchmarking Gender and Political Bias in Large Language Models

Jinrui Yang, Xudong Han, Timothy Baldwin

We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-cal…

cs.CL2025

ParlAI Vote: A Web Platform for Analyzing Gender and Political Bias in Large Language Models

Wenjie Lin, Hange Liu, Yingying Zhuang +5

We present ParlAI Vote, an interactive web platform for exploring European Parliament debates and votes, and for testing LLMs on vote prediction and bias analysis. This web system…

cs.CL2025

RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises

Zenan Zhai, Hao Li, Xudong Han +4

Recent advances in large language models (LLMs) have shown that they can answer questions requiring complex reasoning. However, their ability to identify and respond to text contai…

cs.CL2024

ToolGen: Unified Tool Retrieval and Calling via Generation

Renxi Wang, Xudong Han, Lei Ji +3

As large language models (LLMs) advance, their inability to autonomously execute tasks by directly interacting with external tools remains a critical limitation. Traditional method…

cs.CL2024

Against The Achilles' Heel: A Survey on Red Teaming for Generative Models

Lizhi Lin, Honglin Mu, Zenan Zhai +9

Generative models are rapidly gaining popularity and being integrated into everyday applications, raising concerns over their safe use as various vulnerabilities are exposed. In li…

cs.CL2024

A Chinese Dataset for Evaluating the Safeguards in Large Language Models

Yuxia Wang, Zenan Zhai, Haonan Li +6

Many studies have demonstrated that large language models (LLMs) can produce harmful responses, exposing users to unexpected risks when LLMs are deployed. Previous studies have pro…