2 papers
cs.AI2026
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
Yuzhe Wang, Yaochen Zhu, Jundong Li
As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground their reasoning in causality rather…
cs.CL2026
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
Xiaoyuan Li, Yuzhe Wang, Moxin Li +6
Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under qu…