3 papers
cs.CV2026
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, Lucas Goncalves +7
Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the…
cs.CV2024
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
Wenliang Zhong, Wenyi Wu, Qi Li +6
Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapt…
cs.CV2023
MIVC: Multiple Instance Visual Component for Visual-Language Models
Wenyi Wu, Qi Li, Wenliang Zhong +1
Vision-language models have been widely explored across a wide range of tasks and achieve satisfactory performance. However, it's under-explored how to consolidate entity understan…