13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CV2026
Narrative Aligned Long Form Video Question Answering
Rahul Jain, Keval Doshi, Burak Uzkent +1
Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and…
cs.RO2022★ 13 cited
Learning Neuro-symbolic Programs for Language Guided Robot Manipulation
Namasivayam Kalithasan, Himanshu Singh, Vishal Bindal +5
Given a natural language instruction and an input scene, our goal is to train a model to output a manipulation program that can be executed by the robot. Prior approaches for this…