2 papers
cs.CV2026
DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime
Julian Lorenz, Vladyslav Kovganko, Elias Kohout +3
Scene Graph Generation (SGG) aims to extract a detailed graph structure from an image, a representation that holds significant promise as a robust intermediate step for complex dow…
cs.CL2025
Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models
Christian Jaumann, Annemarie Friedrich, Rainer Lienhart
This paper describes our system for the SciVQA 2025 Shared Task on Scientific Visual Question Answering. Our system employs an ensemble of two Multimodal Large Language Models and…