1 citations · 1 across the 2 of their papers we have counts for
4 papers
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
Sunqi Fan, Lingshan Chen, Runqi Yin +4
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to th…
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Sunqi Fan, Qingle Liu, Runqi Yin +2
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Quest…
PhoneBuddy: Training Open Models for Agentic Phone Use
Zhengyang Tang, Xin Lai, Pengyuan Lyu +23
Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matter…
Agentic Keyframe Search for Video Question Answering
Sunqi Fan, Meng-Hao Guo, Shuojin Yang
Video question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards ach…