5 papers
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
Jongwoo Park, Kanchana Ranasinghe, Kumara Kahatapitiya +3
Long-form videos that span across wide temporal intervals are highly information redundant and contain multiple distinct events or entities that are often loosely related. Therefor…
Understanding Long Videos with Multimodal Language Models
Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya +1
Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world kn…
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Xiang Li, Cristina Mata, Jongwoo Park +8
Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for…
Language Repository for Long Video Understanding
Kumara Kahatapitiya, Kanchana Ranasinghe, Jongwoo Park +1
Language has become a prominent modality in computer vision with the rise of LLMs. Despite supporting long context-lengths, their effectiveness in handling long-term information gr…
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
AJ Piergiovanni, Dahun Kim, Michael S. Ryoo +2
Generating automatic dense captions for videos that accurately describe their contents remains a challenging area of research. Most current models require processing the entire vid…