Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
Roberto Amoroso, Gengyuan Zhang, Rajat Koner +3
Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on cont…
cs.CV2024
LookupViT: Compressing visual information to a limited number of tokens
Rajat Koner, Gagan Jain, Prateek Jain +2
Vision Transformers (ViT) have emerged as the de-facto choice for numerous industry grade vision solutions. But their inference cost can be prohibitive for many settings, as they c…