3 papers
cs.CL2025
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
Zhoutong Ye, Mingze Sun, Huan-ang Gao +9
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by th…
cs.HC2025
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
Yiwen Yin, Zhian Hu, Xiaoxi Xu +4
Measuring GUI task difficulty is crucial for user behavior analysis and agent capability evaluation. Yet, existing benchmarks typically quantify difficulty based on motor actions (…
cs.HC2025
TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
Minghao Tu, Chun Yu, Xiyuan Shen +3
Text boxes serve as portals to diverse functionalities in today's smartphone applications. However, when it comes to specific functionalities, users always need to navigate through…