2 papers
cs.CV2026
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
Sakib Reza, Gauri Jagatap, Mohsen Moghaddam +2
Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We introduce Point-Supervised Onli…
cs.CV2025
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
Sakib Reza, Xiyun Song, Heather Yu +3
Integrating vision models into large language models (LLMs) has sparked significant interest in creating vision-language foundation models, especially for video understanding. Rece…