9 citations · 10 across the 8 of their papers we have counts for
4 papers · 1 filter
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
Sameep Vani, Shreyas Jena, Maitreya Patel +3
While Video Large Language Models (Video-LLMs) have demonstrated remarkable performance across general video understanding benchmarks-particularly in video captioning and descripti…
Integrating Knowledge and Reasoning in Image Understanding
Somak Aditya, Yezhou Yang, Chitta Baral
Deep learning based data-driven approaches have been successfully applied in various image understanding applications ranging from object recognition, semantic segmentation to visu…
Spatial Knowledge Distillation to aid Visual Reasoning
Somak Aditya, Rudra Saha, Yezhou Yang +1
For tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense)…
Explicit Reasoning over End-to-End Neural Architectures for Visual Question Answering
Somak Aditya, Yezhou Yang, Chitta Baral
Many vision and language tasks require commonsense reasoning beyond data-driven image and natural language processing. Here we adopt Visual Question Answering (VQA) as an example t…