2 papers
cs.CV2026
Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning
Zhiyue Liu, Wenkai Zhou, Jian Qin +1
Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from t…
cs.CV2025
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
Zhiyue Liu, Sihang Liu, Jinyuan Liu +1
Knowledge-based visual question answering (KB-VQA) requires a model to understand images and utilize external knowledge to provide accurate answers. Existing approaches often direc…