9 citations · 11 across the 5 of their papers we have counts for
5 papers
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
Yichen Gong, Zhuohan Cai, Sunhao Dai +4
Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of real-world mobile usage. To th…
PC: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
Yue Duan, Zhangxuan Gu, Zhenzhe Ying +3
In the realm of cross-modal retrieval, seamlessly integrating diverse modalities within multimedia remains a formidable challenge, especially given the complexities introduced by n…
E-ANT: A Large-Scale Dataset for Efficient Automatic GUI NavigaTion
Ke Wang, Tianyu Xia, Zhangxuan Gu +5
Online GUI navigation on mobile devices has driven a lot of attention recent years since it contributes to many real-world applications. With the rapid development of large languag…
Backpropagation Path Search On Adversarial Transferability
Zhuoer Xu, Zhangxuan Gu, Jianping Zhang +3
Deep neural networks are vulnerable to adversarial examples, dictating the imperativeness to test the model's robustness before deployment. Transfer-based attackers craft adversari…
Mobile User Interface Element Detection Via Adaptively Prompt Tuning
Zhangxuan Gu, Zhuoer Xu, Haoxing Chen +3
Recent object detection approaches rely on pretrained vision-language models for image-text alignment. However, they fail to detect the Mobile User Interface (MUI) element since it…