4 papers
Resolution as a Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification
Zanwu Liu, Chao Yuan, Bo Li +2
Cross-resolution person re-identification (CR-ReID) remains challenging in practical surveillance, where camera quality and capture distance lead to substantial resolution gaps bet…
CPCL: Cross-Modal Prototypical Contrastive Learning for Weakly Supervised Text-based Person Retrieval
Xinpeng Zhao, Yanwei Zheng, Chuanlin Lan +4
Weakly supervised text-based person retrieval seeks to retrieve images of a target person using textual descriptions, without relying on identity annotations and is more challengin…
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Yifan Zhong, Fengshuo Bai, Shaofei Cai +11
The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence…
Diffusion-based Hierarchical Negative Sampling for Multimodal Knowledge Graph Completion
Guanglin Niu, Xiaowei Zhang
Multimodal Knowledge Graph Completion (MMKGC) aims to address the critical issue of missing knowledge in multimodal knowledge graphs (MMKGs) for their better applications. However,…