2 papers
cs.CV2025
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
Bulat Khaertdinov, Mirela Popa, Nava Tintarev
Large vision-language models (VLMs) enable intuitive visual search using natural language queries. However, improving their performance often requires fine-tuning and scaling to la…
cs.CV2024
Does SpatioTemporal information benefit Two video summarization benchmarks?
Aashutosh Ganesh, Mirela Popa, Daan Odijk +1
An important aspect of summarizing videos is understanding the temporal context behind each part of the video to grasp what is and is not important. Video summarization models have…