4 papers
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
Sakib Reza, Gauri Jagatap, Mohsen Moghaddam +2
Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We introduce Point-Supervised Onli…
psiUnity: A Platform for Multimodal Data-Driven XR
Akhil Ajikumar, Sahil Mayenkar, Steven Yoo +2
Extended reality (XR) research increasingly relies on the ability to stream and synchronize multimodal data between headsets and immersive applications for data-driven interaction…
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Sakib Reza, Xiyun Song, Heather Yu +3
Integrating vision models into large language models (LLMs) has sparked significant interest in creating vision-language foundation models, especially for video understanding. Rece…
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
Sakib Reza, Yuexi Zhang, Mohsen Moghaddam +1
Online video understanding often relies on individual frames, leading to frame-by-frame predictions. Recent advancements such as Online Temporal Action Localization (OnTAL), extend…