9 papers
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
Zhou Liu, Ligang Huang, Zeli Su +5
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying whi…
Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations
Weijie Liang, Yuanfeng Song, Xing Chen +3
Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representati…
Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?
Jing Ye, Yiwen Duan, Yonghong Yu +3
SQL is central to enterprise data engineering, yet generating fully correct SQL code in a single attempt remains difficult, even for experienced developers and advanced text-to-SQL…
Embedding Software Intent: Lightweight Java Module Recovery
Yirui He, Yuqi Huai, Xingyu Chen +1
As an increasing number of software systems reach unprecedented scale, relying solely on code-level abstractions is becoming impractical. While architectural abstractions offer a m…
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
Zhenghao Zhu, Chuxue Cao, Sirui Han +4
In medical data analysis, extracting deep insights from complex, multi-modal datasets is essential for improving patient care, increasing diagnostic accuracy, and optimizing health…
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
Zhou Liu, Zhaoyang Han, Guochen Yan +5
Data governance ensures data quality, security, and compliance through policies and standards, a critical foundation for scaling modern AI development. Recently, large language mod…