154 citations · 308 across the 38 of their papers we have counts for
29 papers · 1 filter
A Sanity Check on Composed Image Retrieval
Yikun Liu, Jiangchao Yao, Weidi Xie +1
Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the desired modification. Despi…
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Yikun Liu, Yuan Liu, Haicheng Wang +6
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
Yuhuan Yang, Xianwei Zhuang, Yuxuan Cai +6
Recent approaches for segmentation have leveraged pretrained generative models as feature extractors, treating segmentation as a downstream adaptation task via indirect feature ret…
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Yikun Liu, Yuan Liu, Shangzhe Di +8
Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
Zichen Wen, Boxue Yang, Shuang Chen +30
We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent…
SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
Zhenjie Mao, Yuhuan Yang, Chaofan Ma +4
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and…