Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
cs.AI2026
HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
Runchuan Zhu, Hongbin Lai, Bowen Jiang +4
Reflection is a powerful mechanism for LLM reasoning, yet its effectiveness hinges on accurately attributing failures to specific reasoning steps, a capability that current models…