3 papers
cs.CV2025
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
Benjamin Reichman, Constantin Patsch, Jack Truxal +2
In outside knowledge visual question answering (OK-VQA), the model must identify relevant visual information within an image and incorporate external knowledge to accurately respon…
cs.CV2025
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
Constantin Patsch, Marsil Zakour, Yuankai Wu +1
In this report, we address the task of online mistake detection, which is vital in domains like industrial automation and education, where real-time video analysis allows human ope…
cs.CV2025
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
David Gastager, Ghazal Ghazaei, Constantin Patsch
Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and co…