2 papers
cs.CV2024
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed +2
Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly…
cs.CV2023
Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding
Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed +3
While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length…