8 papers
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
Yufeng Zhong, Lei Chen, Xuanle Zhao +7
The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology, which extracts information f…
UItron: Foundational GUI Agent with Advanced Perception and Planning
Zhixiong Zeng, Jing Huang, Liming Zheng +7
GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accel…
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
Wenkang Han, Wang Lin, Yiyun Zhou +4
Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when…
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
Wenkang Han, Zhixiong Zeng, Jing Huang +7
Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…
Contrastive Cross-Course Knowledge Tracing via Concept Graph Guided Knowledge Transfer
Wenkang Han, Wang Lin, Liya Hu +6
Knowledge tracing (KT) aims to predict learners' future performance based on historical learning interactions. However, existing KT models predominantly focus on data from a single…
ScaleTrack: Scaling and back-tracking Automated GUI Agents
Jing Huang, Zhixiong Zeng, Wenkang Han +5
Automated GUI agents aims to facilitate user interaction by automatically performing complex tasks in digital environments, such as web, mobile, desktop devices. It receives textua…