124 citations · 171 across the 86 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Jiaxu Qian, Chendong Wang, Yifan Yang +16
Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real w…
cs.CV2024
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
Yuanmin Tang, Xiaoting Qin, Jue Zhang +7
Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user…
cs.CV2024
Sharingan: Extract User Action Sequence from Desktop Recordings
Yanting Chen, Yi Ren, Xiaoting Qin +7
Video recordings of user activities, particularly desktop recordings, offer a rich source of data for understanding user behaviors and automating processes. However, despite advanc…