collaborators

6 papers

cs.CV2026

AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

Xinze Liu, Lei Yang, Dayan Wu +7

Multi-vector representations have emerged as an effective paradigm for multimodal retrieval, representing each sample with multiple complementary embeddings to capture fine-grained…

cs.CV2026

Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing

Hengjie Zhu, Dayan Wu, Zihao Zhang +5

Supervised proxy-based deep cross-modal hashing has become the dominant paradigm for large-scale retrieval. However, prevalent methods model class proxies as deterministic points i…

cs.CV2026

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

Hengjie Zhu, Dayan Wu, Zihao Zhang +6

Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target mo…

cs.CV2026

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

Lei Yang, Xinze Liu, Dayan Wu +7

Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still lack reliable object-level…

cs.CV2026

Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token

Xinze Liu, Ding Wang, Dayan Wu +4

Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing is appealing, but most existi…

cs.CV2025

DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving

Kaiwen Cai, Xinze Liu, Xia Zhou +7

The generation of realistic LiDAR point clouds plays a crucial role in the development and evaluation of autonomous driving systems. Although recent methods for 3D LiDAR point clou…