11 papers · 1 filter
Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models
Zili Zhang, Yilin Wang, Heng Wang +2
Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctly store each individual hop,…
The Deliberative Illusion: Diagnosing Factual Attrition and Stance Homogenization in Multi-Agent LLM Deliberation
Herun Wan, Jiaying Wu, Minnan Luo +4
Multi-agent LLM systems often treat consensus as evidence of successful interaction. For deliberative problems, however, reliability depends on whether agents preserve the facts an…
Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
Shiyan Zheng, Herun Wan, Minnan Luo +1
While existing social bot detectors perform well on benchmarks, their robustness across diverse real-world scenarios remains limited due to unclear ground truth and varied misleadi…
The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence
Herun Wan, Jiaying Wu, Minnan Luo +3
To reliably assist human decision-making, LLMs must maintain factual internal beliefs against misleading injections. While current models resist explicit misinformation, we uncover…
DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
Herun Wan, Jiaying Wu, Minnan Luo +3
Generating textual rationales from large vision-language models (LVLMs) to support trainable multimodal misinformation detectors has emerged as a promising paradigm. However, its e…
GuessBench: Sensemaking Multimodal Creativity in the Wild
Zifeng Zhu, Shangbin Feng, Herun Wan +3
We propose GuessBench, a novel benchmark that evaluates Vision Language Models (VLMs) on modeling the pervasive, noisy, and pluralistic human creativity. GuessBench sources data fr…