12 papers
Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization
Yitian Shi, Di Wen, Zhengqi Han +6
Learning from human video demonstrations remains challenging due to noisy hand-object interactions, unseen objects with partial observation, and cross-embodiment discrepancy. To ad…
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
David J. Lerch, Livien Majer, Zeyun Zhong +3
Robust self-supervised learning of multi-modal video representations is critical for real-world applications such as driver distraction detection, where multiple sensors provide co…
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
David J. Lerch, Sarath Mulugurthi, Manuel Martin +2
Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on general datasets and struggle…
AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
Omar Moured, Jiaming Zhang, M. Saquib Sarfraz +1
Chart summarization is a crucial task for blind and visually impaired individuals as it is their primary means of accessing and interpreting graphical data. Crafting high-quality d…
Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems
Qihao Yuan, Kailai Li, Jiaming Zhang
3D visual grounding (3DVG) aims to locate objects in a 3D scene with natural language descriptions. Supervised methods have achieved decent accuracy, but have a closed vocabulary a…
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
Alexander Vogel, Omar Moured, Yufan Chen +2
Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understandi…