3 papers
cs.CV2024
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
Roberto Amoroso, Gengyuan Zhang, Rajat Koner +3
Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on cont…
cs.CV2024
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
Federico Betti, Lorenzo Baraldi, Rita Cucchiara +1
Diffusion models have significantly advanced generative AI, but they encounter difficulties when generating complex combinations of multiple objects. As the final result heavily de…
cs.CV2024
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi +1
Fine-tuning image captioning models with hand-crafted rewards like the CIDEr metric has been a classical strategy for promoting caption quality at the sequence level. This approach…