activity
20172022
most citedDiversity-aware Multi-Video Summarization

62 citations · 89 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV2022

GraphMapper: Efficient Visual Navigation by Scene Graph Generation

Zachary Seymour, Niluthpol Chowdhury Mithun, Han-Pang Chiu +2

Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, u…

cs.RO20214 cited

SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments

Muhammad Zubair Irshad, Niluthpol Chowdhury Mithun, Zachary Seymour +3

This paper presents a novel approach for the Vision-and-Language Navigation (VLN) task in continuous 3D environments, which requires an autonomous agent to follow natural language…

cs.CV20212 cited

MaAST: Map Attention with Semantic Transformersfor Efficient Visual Navigation

Zachary Seymour, Kowshik Thopalli, Niluthpol Mithun +3

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potentia…

cs.CV202021 cited

RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization

Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu +2

We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (…

cs.CV2019

Weakly Supervised Video Moment Retrieval From Text Queries

Niluthpol Chowdhury Mithun, Sujoy Paul, Amit K. Roy-Chowdhury

There have been a few recent methods proposed in text to video moment retrieval using natural language queries, but requiring full supervision during training. However, acquiring a…

cs.MM2018

Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval

Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis +1

Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great pr…