activity
20242026
collaborators

10 papers

cs.CV2026

AB: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning

Yiyun Zhou, Zhonghua Jiang, Wenkang Han +4

Efficient transfer learning methods for large-scale vision-language models (, CLIP) enable strong few-shot transfer, yet existing adaptation methods follow a fixed fine-tunin…

cs.CV2026

OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models

Yufeng Zhong, Lei Chen, Xuanle Zhao +7

The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology, which extracts information f…

cs.CL2025

UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions

Wenkang Han, Zhixiong Zeng, Jing Huang +7

Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…

cs.CV2025

UItron: Foundational GUI Agent with Advanced Perception and Planning

Zhixiong Zeng, Jing Huang, Liming Zheng +7

GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accel…

cs.CV2025

Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration

Wenkang Han, Wang Lin, Yiyun Zhou +4

Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when…

cs.LG2025

DKT2: Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data

Yiyun Zhou, Wenkang Han, Jingyuan Chen

Knowledge Tracing (KT) is a fundamental component of Intelligent Tutoring Systems (ITS), enabling the modeling of students' knowledge states to predict future performance. The intr…