2 citations · 2 across the 3 of their papers we have counts for
4 papers · 1 filter
Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
Jiasheng Li, Zhong Ji, Yan Zhang +1
Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly impro…
TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering
Zhong Ji, Keqi Jin, Yan Zhang +1
Long-document multimodal question answering requires more than retrieving relevant chunks from a large document. Different queries require different evidence behavior. Existing mul…
Underlying Semantic Diffusion for Effective and Efficient In-Context Learning
Zhong Ji, Weilong Cao, Yan Zhang +3
Diffusion models has emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlyin…
Asymmetric Cross-Scale Alignment for Text-Based Person Search
Zhong Ji, Junhua Hu, Deyin Liu +2
Text-based person search (TBPS) is of significant importance in intelligent surveillance, which aims to retrieve pedestrian images with high semantic relevance to a given text desc…