Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
Longrong Yang, Zhixiong Zeng, Yufeng Zhong +7
Multimodal large language models are evolving toward multimodal agents capable of proactively executing tasks. Most agent research focuses on GUI or embodied scenarios, which corre…
cs.CV2025
UItron: Foundational GUI Agent with Advanced Perception and Planning
Zhixiong Zeng, Jing Huang, Liming Zheng +7
GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accel…
cs.CV2025
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
Yufeng Zhong, Zhixiong Zeng, Lei Chen +5
Optical Character Recognition (OCR) for mathematical formula is essential for the intelligent analysis of scientific literature. However, both task-specific and general vision-lang…