activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

Zihao Zhang, Dayan Wu, Xinze Liu +6

Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existi…

cs.CV2026

AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

Xinze Liu, Lei Yang, Dayan Wu +7

Multi-vector representations have emerged as an effective paradigm for multimodal retrieval, representing each sample with multiple complementary embeddings to capture fine-grained…

cs.CV2026

Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing

Hengjie Zhu, Dayan Wu, Zihao Zhang +5

Supervised proxy-based deep cross-modal hashing has become the dominant paradigm for large-scale retrieval. However, prevalent methods model class proxies as deterministic points i…

cs.CV2026

Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token

Xinze Liu, Ding Wang, Dayan Wu +4

Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing is appealing, but most existi…

cs.CV2025

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning

Zexian Yang, Dian Li, Dayan Wu +2

Despite significant advancements in multimodal reasoning tasks, existing Large Vision-Language Models (LVLMs) are prone to producing visually ungrounded responses when interpreting…

cs.CV2024

RealEra: Semantic-level Concept Erasure via Neighbor-Concept Mining

Yufan Liu, Jinyang An, Wanqian Zhang +5

The remarkable development of text-to-image generation models has raised notable security concerns, such as the infringement of portrait rights and the generation of inappropriate…