4 papers
BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs
Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos +2
LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and…
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
Bin Wu, Arastun Mammadli, Xiaoyu Zhang +1
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unli…
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
Xiao Fu, Hossein A. Rahmani, Bin Wu +3
Personalised text generation is essential for user-centric information systems, yet most evaluation methods overlook the individuality of users. We introduce \textbf{PREF}, a \text…
Instruction Tuning With Loss Over Instructions
Zhengyan Shi, Adam X. Yang, Bin Wu +3
Instruction tuning plays a crucial role in shaping the outputs of language models (LMs) to desired styles. In this work, we propose a simple yet effective method, Instruction Model…