3 papers
cs.CV2025
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
Ziyi Wang, Zhi Gao, Jin Chen +3
Domain generalization (DG) aims to learn a model from source domains and apply it to unseen target domains with out-of-distribution data. Owing to CLIP's strong ability to encode s…
cs.CV2025
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
Yongqi Wang, Xinxiao Wu, Shuo Yang +1
Open-vocabulary video visual relationship detection aims to expand video visual relationship detection beyond annotated categories by detecting unseen relationships between both se…
cs.CV2024
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
Yayun Qi, Hongxi Li, Yiqi Song +2
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelli…