activity
20232026
most citedHierarchical Side-Tuning for Vision Transformers

3 citations · 3 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Rui Tang, Wentao Yang, Peirong Zhang +4

Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for…

cs.CV2026

iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

Xuezhi Cui, Dongbo Zhou, Wang Guo +8

Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning iso…

cs.CV2026

SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision

Zejian Kang, Xuanyang Xu, Wentao Yang +6

Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to se…

cs.CV2026

SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space

Zejian Kang, Kai Zheng, Yuanchen Fei +3

Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. Ho…

cs.CV2024

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Jiaxin Zhang, Wentao Yang, Songxuan Lai +2

Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layout…

cs.CV2023★ 3 cited

Hierarchical Side-Tuning for Vision Transformers

Weifeng Lin, Ziheng Wu, Wentao Yang +3

Fine-tuning pre-trained Vision Transformers (ViTs) has showcased significant promise in enhancing visual recognition tasks. Yet, the demand for individualized and comprehensive fin…