Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
Haibin He, Qihuang Zhong, Juhua Liu +3
Video text-based visual question answering (Video TextVQA) task aims to answer questions about videos by leveraging the visual text appearing within the videos. This task poses sig…
cs.CV2025
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
Xuzheng Yang, Junzhuo Liu, Peng Wang +3
Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image groun…
cs.CV2024
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
Junzhuo Liu, Xuzheng Yang, Weiwei Li +1
Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-i…