2 papers
cs.AI2026
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
Zhe Yu, Wenpeng Xing, Yunzhao Wei +4
Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly…
cs.AI2026
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context
Zhe Yu, Wenpeng Xing, Yunzhao Wei +4
Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually govern…