2 papers
cs.CV2026
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Martin Q. Ma, Willis Guo, Aditya Agrawal +4
Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…
cs.CL2024
KaPQA: Knowledge-Augmented Product Question-Answering
Swetha Eppalapally, Daksh Dangi, Chaithra Bhat +8
Question-answering for domain-specific applications has recently attracted much interest due to the latest advancements in large language models (LLMs). However, accurately assessi…