From the 2 of 6 linked papers with an AI index.
6 papers
HyperGS: Fast and Generalizable Gaussian Video Representation
Fatimah Zohra, Chen Zhao, Shuming Liu +2
HyperGS is a feedforward model that predicts Gaussian video representations directly from input video in a single forward pass, achieving orders-of-magnitude faster encoding and be…
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP
Fatimah Zohra, Chen Zhao, Shuming Liu +1
The paper replaces the softmax in CLIP's visual self‑attention with the α‑entmax transform to create sparse attention, which reduces noise from irrelevant tokens and improves dense…
Mindstorms in Natural Language-Based Societies of Mind
Mingchen Zhuge, Haozhe Liu, Francesco Faccio +23
Both Minsky's "society of mind" and Schmidhuber's "learning to think" inspire diverse societies of large multimodal neural networks (NNs) that solve problems by interviewing each o…
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
Shuming Liu, Chen Zhao, Tianqi Xu +1
Large video-language models (VLMs) have demonstrated promising progress in various video understanding tasks. However, their effectiveness in long-form video analysis is constraine…
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
Chen-Lin Zhang, Lin Sui, Shuming Liu +3
Temporal localization in untrimmed videos, which aims to identify specific timestamps, is crucial for video understanding but remains challenging. This task encompasses several sub…
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra +10
Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…