4 papers · 1 filter
MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval
Zihao Zhang, Dayan Wu, Xinze Liu +6
Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existi…
FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding
Hengjie Zhu, Dayan Wu, Zihao Zhang +6
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target mo…
Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
Lei Yang, Xinze Liu, Dayan Wu +7
Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still lack reliable object-level…
Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token
Xinze Liu, Ding Wang, Dayan Wu +4
Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing is appealing, but most existi…