4 papers
MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
Zhichao Yang, Yuanze Hu, Haojie Hao +7
Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, mul…
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading
Yuanze Hu, Gen Li, Yuqin Lan +5
Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measurement reading. In this paper, we…
Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
Zhichao Yang, Zhaoxin Fan, Gen Li +6
Structured, procedural reasoning is essential for Large Language Models (LLMs), especially in mathematics. While post-training methods have improved LLM performance, they still fal…
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
Yuanze Hu, Zhaoxin Fan, Xinyu Wang +8
Lightweight Vision-Language Models (VLMs) are indispensable for resource-constrained applications. The prevailing approach to aligning vision and language models involves freezing…