2 papers
cs.CL2026
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
Tianyi Huang, Nathan Huang, Justin Tang +2
Large language models (LLMs) are now widely used as judges, yet their decisions can change under presentation choices that should be irrelevant. We study one such source of instabi…
cs.CL2025
Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data
Tianyi Huang, Elsa Fan
From disinformation spread by AI chatbots to AI recommendations that inadvertently reinforce stereotypes, textual bias poses a significant challenge to the trustworthiness of large…