Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, Lucas Goncalves +7
Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the…
cs.CV2024
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
Wenliang Zhong, Wenyi Wu, Qi Li +6
Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapt…