4 papers
Improving vision-language alignment with graph spiking hybrid Networks
Siyu Zhang, Wenzhe Liu, Yeming Chen +3
To bridge the semantic gap between vision and language (VL), it is necessary to develop a good alignment strategy, which includes handling semantic diversity, abstract representati…
Superpixel Semantics Representation and Pre-training for Vision-Language Task
Siyu Zhang, Yeming Chen, Yaoru Sun +4
The key to integrating visual language tasks is to establish a good alignment strategy. Recently, visual semantic representation has achieved fine-grained visual understanding by d…
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
Haoran Wang, Zeshen Tang, Leya Yang +4
Goal-conditioned hierarchical reinforcement learning (HRL) presents a promising approach for enabling effective exploration in complex, long-horizon reinforcement learning (RL) tas…
Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning
Yeming Chen, Siyu Zhang, Yaoru Sun +2
With the success of self-supervised learning, multimodal foundation models have rapidly adapted a wide range of downstream tasks driven by vision and language (VL) pretraining. Sta…