4 papers
Enhancing Web Agents with a Hierarchical Memory Tree
Yunteng Tan, Zhi Gao, Xinxiao Wu
Large language model-based web agents have shown strong potential in automating web interactions through advanced reasoning and instruction following. While retrieval-based memory…
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
Pengxiang Li, Zechen Hu, Zirui Shang +15
Vision-language model (VLM) based GUI agents show promise for automating complex desktop and mobile tasks, but face significant challenges in applying reinforcement learning (RL):…
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
Ziyi Wang, Zhi Gao, Jin Chen +3
Domain generalization (DG) aims to learn a model from source domains and apply it to unseen target domains with out-of-distribution data. Owing to CLIP's strong ability to encode s…
VUDG: A Dataset for Video Understanding Domain Generalization
Ziyi Wang, Zhi Gao, Boxuan Yu +5
Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, existin…