2 papers
cs.CV2025
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
Miaosen Zhang, Ziqiang Xu, Jialiang Zhu +8
With the development of multimodal reasoning models, Computer Use Agents (CUAs), akin to Jarvis from \textit{"Iron Man"}, are becoming a reality. GUI grounding is a core component…
cs.CV2025
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
Ziqiang Xu, Qi Dai, Tian Xie +5
Video understanding is inherently intention-driven-humans naturally focus on relevant frames based on their goals. Recent advancements in multimodal large language models (MLLMs) h…