33 papers
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Kawai Chung, Chunkit Chan, Yauwai Yim +12
The paper introduces MultivationBench, a benchmark that tests multimodal large language models on their ability to reason about evolving human motivations across sequential visual…
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
Jiaxin Bai, Yue Guo, Yifei Dong +13
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
Tian Zheng, Kai-Tai Hsu
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM res…
Target-Aware Linear Regression Under Distribution Shift
Zhewen Hou, Tian Zheng
Distribution shift between training and deployment is a pervasive challenge for modern AI systems. In many cases, the target marginals of covariates and response are known or speci…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
Yueming Wang, Tianshi Zheng, Jiaxin Bai +3
Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
Qiao Xiao, Haochen Shi, Yisen Gao +9
Large language model (LLM) agents increasingly rely on agent harnesses that manage context, tools, and multi-turn execution, making tools a central interface for acting in realisti…