3 papers
cs.CV2026
Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement
Bo Ma
Vision--language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to all…
cs.IR2025
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
Tianyuan Li, Lei Wang, Ahtamjan Ahmat +4
Generative cross-modal retrieval, which treats retrieval as a generation task, has emerged as a promising direction with the rise of Multimodal Large Language Models (MLLMs). In th…
cs.CV2025
TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement
Jibai Lin, Bo Ma, Yating Yang +6
Subject-driven image generation (SDIG) aims to manipulate specific subjects within images while adhering to textual instructions, a task crucial for advancing text-to-image diffusi…