4 papers
EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
Runze Li, Yuwen Zhai, Bo Xu +5
Contemporary GUI agents, while increasingly capable due to advances in Large Vision-Language Models (VLMs), often operate with a critical limitation: they treat each task in isolat…
GUIDE: Interpretable GUI Agent Evaluation via Hierarchical Diagnosis
Yuwen Zhai, Runze Li, Liang Wang +6
Evaluating GUI agents presents a distinct challenge: trajectories are long, visually grounded, and open-ended, yet evaluation must be both accurate and interpretable. Existing appr…
LOST: Low-rank and Sparse Pre-training for Large Language Models
Jiaxi Li, Lu Yin, Li Shen +6
While large language models (LLMs) have achieved remarkable performance across a wide range of tasks, their massive scale incurs prohibitive computational and memory costs for pre-…
CLIP Brings Better Features to Visual Aesthetics Learners
Liwu Xu, Jinjin Xu, Yuzhe Yang +3
Image Aesthetics Assessment (IAA) is a challenging task due to its subjective nature and expensive manual annotations. Recent large-scale vision-language models, such as Contrastiv…