2 papers
cs.AI2026
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
Zixing Lin, Jiale Wang, Gee Wah Ng +4
Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q T…
cs.CV2026
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
Jiale Wang, Gee Wah Ng, Lee Onn Mak +3
This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM)…