4 papers · 1 filter
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
Benjamin Reichman, Constantin Patsch, Jack Truxal +2
In outside knowledge visual question answering (OK-VQA), the model must identify relevant visual information within an image and incorporate external knowledge to accurately respon…
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
Constantin Patsch, Marsil Zakour, Yuankai Wu +1
In this report, we address the task of online mistake detection, which is vital in domains like industrial automation and education, where real-time video analysis allows human ope…
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
David Gastager, Ghazal Ghazaei, Constantin Patsch
Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and co…
ADL4D: Towards A Contextually Rich Dataset for 4D Activities of Daily Living
Marsil Zakour, Partha Pratim Nath, Ludwig Lohmer +6
Hand-Object Interactions (HOIs) are conditioned on spatial and temporal contexts like surrounding objects, previous actions, and future intents (for example, grasping and handover…