10 citations · 10 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Efficient Multi-modal Long Context Learning for Training-free Adaptation
Zehong Ma, Shiliang Zhang, Longhui Wei +1
Traditional approaches to adapting multi-modal large language models (MLLMs) to new tasks have relied heavily on fine-tuning. This paper introduces Efficient Multi-Modal Long Conte…
cs.CV2025★ 10 cited
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
Zehong Ma, Hao Chen, Wei Zeng +2
Fine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately…
cs.CV2024
OVMR: Open-Vocabulary Recognition with Multi-Modal References
Zehong Ma, Shiliang Zhang, Longhui Wei +1
The challenge of open-vocabulary recognition lies in the model has no clue of new categories it is applied to. Existing works have proposed different methods to embed category cues…