3 papers
cs.CV2026
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
Rui Shao, Ruize Gao, Bin Xie +5
Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-qua…
cs.RO2026
EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models
Mingchen Song, Xiang Deng, Jie Wei +3
Recent advances in unimanual manipulation policies have achieved remarkable success across diverse robotic tasks through abundant training data and well-established model architect…
cs.RO2025
Few-Shot Vision-Language Action-Incremental Policy Learning
Mingchen Song, Xiang Deng, Guoqiang Zhong +5
Recently, Transformer-based robotic manipulation methods utilize multi-view spatial representations and language instructions to learn robot motion trajectories by leveraging numer…