6 papers
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Changyuan Wang, Chubin Zhang, Zhenyu Wu +8
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. How…
Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval
Runhao Li, Xiaoxu Ma, Zhenyu Weng +5
Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual semantic annotation. However, ex…
Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing
Runhao Li, Xiaoxu Ma, Zhenyu Weng +5
Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existi…
UniHash: Unifying Pointwise and Pairwise Hashing Paradigms
Xiaoxu Ma, Runhao Li, Xiangbo Zhang +1
Effective retrieval across both seen and unseen categories is crucial for modern image retrieval systems. Retrieval on seen categories ensures precise recognition of known classes,…
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
Runhao Li, Wenkai Guo, Zhenyu Wu +5
Pre-trained Vision-Language-Action (VLA) models have achieved remarkable success in improving robustness and generalization for end-to-end robotic manipulation. However, these mode…
Mutual Learning for Hashing: Unlocking Strong Hash Functions from Weak Supervision
Xiaoxu Ma, Runhao Li, Zhenyu Weng
Deep hashing has been widely adopted for large-scale image retrieval, with numerous strategies proposed to optimize hash function learning. Pairwise-based methods are effective in…