3 papers
cs.CV2026
ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding
Yuanhao Sun, Huawei Ji, Jiaxin Ding +2
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising objec…
cs.CV2026
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models
Huawei Ji, Yuanhao Sun, Yuan Jin +4
Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs).…
cs.CL2025
AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing
Huawei Ji, Cheng Deng, Bo Xue +6
With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predomin…