3 papers
cs.LG2026
SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis
Zhuohang Fan, Beichen Zhang, Yuanfa Li +4
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-covera…
cs.CV2026
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
Yanpei Gong, Beichen Zhang, Hao Wang +6
Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored video modeling, they largely ov…
cs.CV2025
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
Yufan Zhou, Zhaobo Qi, Lingshuai Lin +5
In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual obser…