5 citations · 5 across the 6 of their papers we have counts for
4 papers · 1 filter
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
Ghazal Kaviani, Ghassan AlRegib
Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases…
Multi-level and Multi-modal Action Anticipation
Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar +1
Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates o…
A Large-scale Benchmark on Geological Fault Delineation Models: Domain Shift, Training Dynamics, Generalizability, Evaluation and Inferential Behavior
Jorge Quesada, Chen Zhou, Prithwijit Chowdhury +5
Machine learning has taken a critical role in seismic interpretation workflows, especially in fault delineation tasks. However, despite the recent proliferation of pretrained model…
Hierarchical and Multimodal Data for Daily Activity Understanding
Ghazal Kaviani, Yavuz Yarici, Seulgi Kim +4
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in…