62 citations · 89 across the 5 of their papers we have counts for
7 papers
GraphMapper: Efficient Visual Navigation by Scene Graph Generation
Zachary Seymour, Niluthpol Chowdhury Mithun, Han-Pang Chiu +2
Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, u…
SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
Muhammad Zubair Irshad, Niluthpol Chowdhury Mithun, Zachary Seymour +3
This paper presents a novel approach for the Vision-and-Language Navigation (VLN) task in continuous 3D environments, which requires an autonomous agent to follow natural language…
MaAST: Map Attention with Semantic Transformersfor Efficient Visual Navigation
Zachary Seymour, Kowshik Thopalli, Niluthpol Mithun +3
Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potentia…
RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization
Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu +2
We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (…
Weakly Supervised Video Moment Retrieval From Text Queries
Niluthpol Chowdhury Mithun, Sujoy Paul, Amit K. Roy-Chowdhury
There have been a few recent methods proposed in text to video moment retrieval using natural language queries, but requiring full supervision during training. However, acquiring a…
Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval
Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis +1
Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great pr…