Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning
Hyeong Kyu Choi, Xiaojin Zhu, Sharon Li
Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their opinions. Yet recent studies reve…
cs.AI2026
Beyond In-Domain Detection: SpikeScore for Cross-Domain Hallucination Detection
Yongxin Deng, Zhen Fang, Sharon Li +1
Hallucination detection is critical for deploying large language models (LLMs) in real-world applications. Existing hallucination detection methods achieve strong performance when…