Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
Dayong Liang, Changmeng Zheng, Zhiyuan Wen +3
Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper…
cs.CV2025
PolySmart @ TRECVid 2024 Video Captioning (VTT)
Jiaxin Wu, Wengyu Zhang, Xiao-Yong Wei +1
In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA…
cs.CV2024
PolySmart @ TRECVid 2024 Medical Video Question Answering
Jiaxin Wu, Yiyang Jiang, Xiao-Yong Wei +1
Video Corpus Visual Answer Localization (VCVAL) includes question-related video retrieval and visual answer localization in the videos. Specifically, we use text-to-text retrieval…