1 citations · 1 across the 10 of their papers we have counts for
4 papers · 1 filter
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
Sosuke Yamao, Natsuki Miyahara, Yuankai Qi +1
In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented…
ProgRoCC: A Progressive Approach to Rough Crowd Counting
Shengqin Jiang, Linfei Li, Haokui Zhang +6
As the number of individuals in a crowd grows, enumeration-based techniques become increasingly infeasible and their estimates increasingly unreliable. We propose instead an estima…
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
Huajie Jiang, Zhengxian Li, Xiaohan Yu +4
Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires c…
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
Zhuo Tao, Liang Li, Qi Chen +5
Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Rec…