5 papers
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Xuan He, Cong Wei, Yuhao Cheng +19
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: b…
Sanity Checks for Agentic Data Science
Zachary T. Rewolinski, Austin V. Zane, Hao Huang +4
Agentic data science (ADS) pipelines have grown rapidly in both capability and adoption, with systems such as OpenAI Codex now able to directly analyze datasets and produce answers…
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
Ke Yang, Zixi Chen, Xuan He +6
Long-term memory is essential for large language model (LLM) agents operating in complex environments, yet existing memory designs are either task-specific and non-transferable, or…
ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation
Yimeng Liu, Misha Sra, Jeevana Priya Inala +1
AI-powered web agents have the potential to automate repetitive tasks, such as form filling, information retrieval, and scheduling, but they struggle to reliably execute these task…
Towards Understanding Graphical Perception in Large Multimodal Models
Kai Zhang, Jianwei Yang, Jeevana Priya Inala +4
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprising…