2 papers
cs.CV2025
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
Haodi Ma, Vyom Pathak, Daisy Zhe Wang
Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often re…
cs.DB2024
Xling: A Learned Filter Framework for Accelerating High-Dimensional Approximate Similarity Join
Yifan Wang, Vyom Pathak, Daisy Zhe Wang
Similarity join finds all pairs of close points within a given distance threshold. Many similarity join methods have been proposed, but they are usually not efficient on high-dimen…