10 papers
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, Sahil Shah +4
We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek…
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
Minkyu Choi, S P Sharan, Harsh Goel +2
Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to…
SSR: A Generic Framework for Text-Aided Map Compression for Localization
Mohammad Omama, Po-han Li, Harsh Goel +6
Mapping is crucial in robotics for localization and downstream decision-making. As robots are deployed in ever-broader settings, the maps they rely on continue to increase in size.…
Pocket RAG: On-Device RAG for First Aid Guidance in Offline Mobile Environment
Dong Ho Kang, Hyunjoon Lee, Hyeonjeong Cha +2
In disaster scenarios or remote areas, first responders often lose network connectivity when providing first aid. In such situations, server-based AI systems fail to provide critic…
ObjectAlign: Neuro-Symbolic Object Consistency Verification and Correction
Mustafa Munir, Harsh Goel, Xiwen Wei +6
Video editing and synthesis often introduce object inconsistencies, such as frame flicker and identity drift that degrade perceptual quality. To address these issues, we introduce…
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
Sahil Shah, S P Sharan, Harsh Goel +5
While vision-language models (VLMs) excel at tasks involving single images or short videos, they still struggle with Long Video Question Answering (LVQA) due to its demand for comp…