79 citations · 79 across the 1 of their papers we have counts for
5 papers
Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
Zeyi Huang, Yuyang Ji, Xiaofang Wang +11
Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited cont…
Human Action Anticipation: A Survey
Bolin Lai, Sam Toyer, Tushar Nagarajan +7
Predicting future human behavior is an increasingly popular topic in computer vision, driven by the interest in applications such as autonomous vehicles, digital assistants and hum…
SF-Net: Single-Frame Supervision for Temporal Action Localization
Fan Ma, Linchao Zhu, Yi Yang +4
In this paper, we study an intermediate form of supervision, i.e., single-frame supervision, for temporal action localization (TAL). To obtain the single-frame supervision, the ann…
Only Time Can Tell: Discovering Temporal Data for Temporal Modeling
Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan +3
Understanding temporal information and how the visual world changes over time is a fundamental ability of intelligent systems. In video understanding, temporal information is at th…
Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
Shengxin Zha, Florian Luisier, Walter Andrews +2
We conduct an in-depth exploration of different strategies for doing event detection in videos using convolutional neural networks (CNNs) trained for image classification. We study…