1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
Wenbin Wang, Yongcheng Jing, Liang Ding +5
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To overcome the limitations of existing methods, this paper shifts away f…
cs.CV2024
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
Xinyu Ma, Xuebo Liu, Derek F. Wong +6
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated tha…