3 papers
cs.CV2026
Online Reasoning Video Object Segmentation
Jinyuan Liu, Yang Wang, Zeyu Zhao +3
Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existi…
cs.CV2025
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
Zhiyue Liu, Sihang Liu, Jinyuan Liu +1
Knowledge-based visual question answering (KB-VQA) requires a model to understand images and utilize external knowledge to provide accurate answers. Existing approaches often direc…
cs.CV2023
Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning
Zhiyue Liu, Jinyuan Liu, Fanrong Ma
Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts…