Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
arXiv:2410.09380 · doi:10.1109/TCSVT.2024.3475510
Abstract
Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model's focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model's ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA.
IEEE Transactions on Circuits and Systems for Video Technology
References in corpus (27)
- LLaMA: Open and Efficient Foundation Language Models
- Learning to Prompt for Vision-Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Is Space-Time Attention All You Need for Video Understanding?
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Scaling Instruction-Finetuned Language Models
- Zero-Shot Text-to-Image Generation
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Visual Instruction Tuning
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
- OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
- Learning Universal Policies via Text-Guided Video Generation
- Self-Chained Image-Language Model for Video Localization and Question Answering
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
- Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- Toward Fairness Through Fair Multi-Exit Framework for Dermatological Disease Diagnosis
- Paxion: Patching Action Knowledge in Video-Language Foundation Models
- Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
- Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer
- Semi-Parametric Video-Grounded Text Generation
- CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
- Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge