2 papers
cs.CL2026
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
Weimin Xiong, Shuhao Gu, Bowen Ye +5
Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarc…
cs.CV2026
Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation
Feifan Song, Yuntian Bo, Haofeng Zhang
Cross-domain few-shot medical image segmentation (CD-FSMIS) requires a model to generalise simultaneously to novel anatomical categories and unseen imaging domains from only a hand…