data scaling 1environment simulation 1mobile gui agents 1multimodal models 1planning and reflection 1reinforcement learning 1
From the 1 of 3 linked papers with an AI index.
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
cs.CV2025
Beyond Intermediate States: Explaining Visual Redundancy through Language
Dingchen Yang, Bowen Cao, Anran Zhang +3
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…
cs.CV2024
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
Dingchen Yang, Bowen Cao, Guang Chen +1
Multi-modal Large Language Models (MLLMs) demonstrate remarkable success across various vision-language tasks. However, they suffer from visual hallucination, where the generated r…