Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
Ghazal Kaviani, Ghassan AlRegib
Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases…
cs.CV2025
Multi-level and Multi-modal Action Anticipation
Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar +1
Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates o…
cs.CV2025
Hierarchical and Multimodal Data for Daily Activity Understanding
Ghazal Kaviani, Yavuz Yarici, Seulgi Kim +4
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in…