Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
arXiv:2410.09380 · doi:10.1109/TCSVT.2024.3475510
Abstract
Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model's focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model's ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA.
IEEE Transactions on Circuits and Systems for Video Technology
References in corpus (14)
- LLaMA: Open and Efficient Foundation Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Scaling Instruction-Finetuned Language Models
- Zero-Shot Text-to-Image Generation
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
- OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
- A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
- Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
- Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer
- Semi-Parametric Video-Grounded Text Generation