QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
arXiv:2601.06566 · doi:10.23919/FUSION59988.2024.10706514
Abstract
This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM) for image-text analysis, and a Large Language Model (LLM) for text analysis. This approach enables integrated analysis of text, images, and video, achieving performance improvements over existing video captioning and Q&A models; all while remaining fully self-contained, adept for on-premises deployment. Experimental results using QCaption demonstrated up to 44.2% and 48.9% improvements in video captioning and Q&A tasks, respectively. Ablation studies were also performed to assess the role of LLM on the fusion on the results. Moreover, the paper proposes and evaluates additional video captioning approaches, benchmarking them against QCaption and existing methodologies. QCaption demonstrate the potential of adopting a model fusion approach in advancing video analytics.
References in corpus (10)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- PaLM: Scaling Language Modeling with Pathways
- Gemini: A Family of Highly Capable Multimodal Models
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- Instruction Tuning with GPT-4
- Does a Technique for Building Multimodal Representation Matter? -- Comparative Analysis
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Deep Model Fusion: A Survey
- A Review of Deep Learning for Video Captioning