5 papers · 1 filter
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
Zhenwei Shao, Mingyang Wang, Weijun Zhang +6
Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges f…
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
Ting Yu, Kunhao Fu, Shuhui Wang +2
Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and s…
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
Ting Yu, Kunhao Fu, Jian Zhang +2
Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-fo…
Learning to Discover Knowledge: A Weakly-Supervised Partial Domain Adaptation Approach
Mengcheng Lan, Min Meng, Jun Yu +1
Domain adaptation has shown appealing performance by leveraging knowledge from a source domain with rich annotations. However, for a specific target task, it is cumbersome to colle…
Boundary Discretization and Reliable Classification Network for Temporal Action Detection
Zhenying Fang, Jun Yu, Richang Hong
Temporal action detection aims to recognize the action category and determine each action instance's starting and ending time in untrimmed videos. The mixed methods have achieved r…