4 papers
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
Ghazal Kaviani, Ghassan AlRegib
Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases…
Multi-level and Multi-modal Action Anticipation
Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar +1
Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates o…
Hierarchical and Multimodal Data for Daily Activity Understanding
Ghazal Kaviani, Yavuz Yarici, Seulgi Kim +4
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in…
Evaluating BM3D and NBNet: A Comprehensive Study of Image Denoising Across Multiple Datasets
Ghazal Kaviani, Reza Marzban, Ghassan AlRegib
This paper investigates image denoising, comparing traditional non-learning-based techniques, represented by Block-Matching 3D (BM3D), with modern learning-based methods, exemplifi…