2 papers
cs.AI2026
Benchmarking World-Model Learning with Environment-Level Queries
Archana Warrier, Dat Nguyen, Michelangelo Naim +8
World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, s…
cs.IR2025
Had enough of experts? Quantitative knowledge retrieval from large language models
David Selby, Kai Spriestersbach, Yuichiro Iwashita +6
Large language models (LLMs) have been extensively studied for their abilities to generate convincing natural language sequences, however their utility for quantitative information…