9 papers
HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Jinhong Deng, Limeng Qiao, Guanglu Wan
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Languag…
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
Yuhang Yang, Jinhong Deng, Wen Li +1
While vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle…
Object-Centric Vision Token Pruning for Vision Language Models
Guangyuan Li, Rongzhen Zhao, Jinhong Deng +2
In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning r…
Mapping Text to Multiplex Graph: Prompt Compression as Lévy Walk-Guided Graph Pruning
Yaxin Gao, Yao Lu, Jinhong Deng +7
Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is often spread across multiple lo…
Deformation-based In-Context Learning for Point Cloud Understanding
Chengxing Lin, Jinhong Deng, Yinjie Lei +1
Recent advances in point cloud In-Context Learning (ICL) have demonstrated strong multitask capabilities. Existing approaches typically adopt a Masked Point Modeling (MPM)-based pa…
Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression
Chenyue Yu, Lingao Xiao, Jinhong Deng +2
Large-scale image datasets are fundamental to deep learning, but their high storage demands pose challenges for deployment in resource-constrained environments. While existing appr…