3 papers
cs.RO2025
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
Wei Zhao, Pengxiang Ding, Min Zhang +4
Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heav…
cs.CV2024
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
Can Cui, Siteng Huang, Wenxuan Song +3
To address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. Howe…
cs.CV2024
Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning
Fengyuan Dai, Siteng Huang, Min Zhang +2
To transfer knowledge from seen attribute-object compositions to recognize unseen ones, recent compositional zero-shot learning (CZSL) methods mainly discuss the optimal classifica…