1 paper
Jason Nguyen, Ameet Rao, Alexander Chang +2
Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step…