24 citations · 35 across the 32 of their papers we have counts for
32 papers
Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Object Category-agnostic Temporal Localization
Yitian Shi, Di Wen, Zhengqi Han +6
Human videos provide a scalable source of demonstrations for robot imitation learning. However, converting them into executable and semantically aligned robot trajectories requires…
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
David J. Lerch, Livien Majer, Zeyun Zhong +3
Robust self-supervised learning of multi-modal video representations is critical for real-world applications such as driver distraction detection, where multiple sensors provide co…
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
David J. Lerch, Sarath Mulugurthi, Manuel Martin +2
Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on general datasets and struggle…
CHAOS: Chart Analysis with Outlier Samples
Omar Moured, Yufan Chen, Ruiping Liu +4
Charts play a critical role in data analysis and visualization, yet real-world applications often present charts with challenging or noisy features. However, "outlier charts" pose…
Scene-agnostic Pose Regression for Visual Localization
Junwei Zheng, Ruiping Liu, Yufan Chen +4
Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes bett…
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
Alexander Vogel, Omar Moured, Yufan Chen +2
Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understandi…