327 citations · 1.5k across the 313 of their papers we have counts for
46 papers · 1 filter
Training Object Permanence in World Models
Haotian Zhang, Fengyuan Yu, Dezhi Luo +28
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun t…
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
Griffin Farrow, Lily Sijia Li, Jack Johnson +4
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading…
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases
Yingqian Wu, Jingcong Liang, Siyuan Wang +4
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research idea…
LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems
Heng Zhou, Lian Zhang, Yutao Fan +5
Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be truste…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
William Bolton, Philip Torr
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by fra…