2 papers
cs.CV2025
Improving vision-language alignment with graph spiking hybrid Networks
Siyu Zhang, Wenzhe Liu, Yeming Chen +3
To bridge the semantic gap between vision and language (VL), it is necessary to develop a good alignment strategy, which includes handling semantic diversity, abstract representati…
cs.CV2024
KNVQA: A Benchmark for evaluation knowledge-based VQA
Sirui Cheng, Siyu Zhang, Jiayi Wu +1
Within the multimodal field, large vision-language models (LVLMs) have made significant progress due to their strong perception and reasoning capabilities in the visual and languag…