6 papers
Agentic Data Environments
Elaine Ang, Chenxi Huang, Georgios Liargkovas +13
Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. The central challenge for agen…
LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake
Haonan Wang, Jiaxiang Liu, Yurong Liu +11
Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In cont…
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
Qiran Zhang, Yuheng Wang, Runde Yang +9
Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models…
TeachMaster: Generative Teaching via Code
Yuheng Wang, Runde Yang, Lin Wu +9
The scalability of high-quality online education is hindered by the high costs and slow cycles of manual content creation. Despite advancements in video generation, current approac…
An approach for systematic decomposition of complex llm tasks
Tianle Zhou, Jiakai Xu, Guanhong Liu +3
Large Language Models (LLMs) suffer from reliability issues on complex tasks, as existing decomposition methods are heuristic and rely on agent or manual decomposition. This work i…
Toward Systems Foundations for Agentic Exploration
Jiakai Xu, Tianle Zhou, Eugene Wu +1
Agentic exploration, letting LLM-powered agents branch, backtrack, and search across many execution paths, demands systems support well beyond today's pass-at-k resets. Our benchma…