10 papers
AB: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
Yiyun Zhou, Zhonghua Jiang, Wenkang Han +4
Efficient transfer learning methods for large-scale vision-language models (, CLIP) enable strong few-shot transfer, yet existing adaptation methods follow a fixed fine-tunin…
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
Yufeng Zhong, Lei Chen, Xuanle Zhao +7
The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology, which extracts information f…
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
Wenkang Han, Zhixiong Zeng, Jing Huang +7
Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…
UItron: Foundational GUI Agent with Advanced Perception and Planning
Zhixiong Zeng, Jing Huang, Liming Zheng +7
GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accel…
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
Wenkang Han, Wang Lin, Yiyun Zhou +4
Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when…
DKT2: Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data
Yiyun Zhou, Wenkang Han, Jingyuan Chen
Knowledge Tracing (KT) is a fundamental component of Intelligent Tutoring Systems (ITS), enabling the modeling of students' knowledge states to predict future performance. The intr…