2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2022
When Can Transformers Ground and Compose: Insights from Compositional Generalization Benchmarks
Ankur Sikarwar, Arkil Patel, Navin Goyal
Humans can reason compositionally whilst grounding language utterances to the real world. Recent benchmarks like ReaSCAN use navigation tasks grounded in a grid world to assess whe…
cs.CV2022★ 2 cited
On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering
Ankur Sikarwar, Gabriel Kreiman
In recent years, multi-modal transformers have shown significant progress in Vision-Language tasks, such as Visual Question Answering (VQA), outperforming previous architectures by…