11 papers
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
Zhou Liu, Ligang Huang, Zeli Su +5
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying whi…
Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations
Weijie Liang, Yuanfeng Song, Xing Chen +3
Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representati…
Self-Evolving Deep Research via Joint Generation and Evaluation
Han Zhu, Chengkun Cai, Yuanfeng Song +3
Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional ques…
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
Zhenghao Zhu, Yuanfeng Song, Xin Chen +5
Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets, we need to perform deep explora…
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
Zhenghao Zhu, Chuxue Cao, Sirui Han +4
In medical data analysis, extracting deep insights from complex, multi-modal datasets is essential for improving patient care, increasing diagnostic accuracy, and optimizing health…
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
Zhou Liu, Zhaoyang Han, Guochen Yan +5
Data governance ensures data quality, security, and compliance through policies and standards, a critical foundation for scaling modern AI development. Recently, large language mod…