2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
Wenliang Zhong, Wenyi Wu, Qi Li +6
Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapt…
cs.CV2023
MIVC: Multiple Instance Visual Component for Visual-Language Models
Wenyi Wu, Qi Li, Wenliang Zhong +1
Vision-language models have been widely explored across a wide range of tasks and achieve satisfactory performance. However, it's under-explored how to consolidate entity understan…
cs.CV2023★ 2 cited
Catalog Phrase Grounding (CPG): Grounding of Product Textual Attributes in Product Images for e-commerce Vision-Language Applications
Wenyi Wu, Karim Bouyarmane, Ismail Tutar
We present Catalog Phrase Grounding (CPG), a model that can associate product textual data (title, brands) into corresponding regions of product images (isolated product region, br…