7 citations · 15 across the 4 of their papers we have counts for
5 papers · 1 filter
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal +2
While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the…
Read My Mind: A Multi-Modal Dataset for Human Belief Prediction
Jiafei Duan, Samson Yu, Nicholas Tan +2
Understanding human intentions is key to enabling effective and efficient human-robot interaction (HRI) in collaborative settings. To enable developments and evaluation of the abil…
Actionet: An Interactive End-To-End Platform For Task-Based Data Collection And Augmentation In 3D Environment
Jiafei Duan, Samson Yu, Hui Li Tan +1
The problem of task planning for artificial agents remains largely unsolved. While there has been increasing interest in data-driven approaches for the study of task planning for a…
6D Pose Estimation with Correlation Fusion
Yi Cheng, Hongyuan Zhu, Ying Sun +6
6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illum…
An End-to-End Network for Generating Social Relationship Graphs
Arushi Goel, Keng Teck Ma, Cheston Tan
Socially-intelligent agents are of growing interest in artificial intelligence. To this end, we need systems that can understand social relationships in diverse social contexts. In…