From the 1 of 6 linked papers with an AI index.
6 papers
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Rui Tang, Wentao Yang, Peirong Zhang +4
The paper introduces a vision-centric framework called SPaTS that uses a single visual token per text instance and reinforcement learning to improve scene text spotting accuracy an…
iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models
Xuezhi Cui, Dongbo Zhou, Wang Guo +8
Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning iso…
SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision
Zejian Kang, Xuanyang Xu, Wentao Yang +6
Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to se…
SparseOIT: Improving Order-Independent Transparency 3DGS via Active Set Method
Wentao Yang, Fanzhen Kong, Zejian Kang +1
3D Gaussian Splatting (3DGS) has received tremendous popularity over the past few years due to its photorealistic visual appearance. However, 3DGS uses volumetric rendering that is…
SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space
Zejian Kang, Kai Zheng, Yuanchen Fei +3
Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. Ho…
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Jiaxin Zhang, Wentao Yang, Songxuan Lai +2
Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layout…