3 papers
cs.CV2025
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
Benjamin Reichman, Constantin Patsch, Jack Truxal +2
In outside knowledge visual question answering (OK-VQA), the model must identify relevant visual information within an image and incorporate external knowledge to accurately respon…
cs.CV2025
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
Constantin Patsch, Marsil Zakour, Yuankai Wu +1
In this report, we address the task of online mistake detection, which is vital in domains like industrial automation and education, where real-time video analysis allows human ope…
cs.CV2024
ADL4D: Towards A Contextually Rich Dataset for 4D Activities of Daily Living
Marsil Zakour, Partha Pratim Nath, Ludwig Lohmer +6
Hand-Object Interactions (HOIs) are conditioned on spatial and temporal contexts like surrounding objects, previous actions, and future intents (for example, grasping and handover…