65 citations · 112 across the 7 of their papers we have counts for
17 papers
Ground then Navigate: Language-guided Navigation in Dynamic Scenes
Kanishk Jain, Varun Chhangani, Amogh Tiwari +2
We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable…
Emotional Prosody Control for Speech Generation
Sarath Sivaprasad, Saiteja Kosgi, Vineet Gandhi
Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selec…
High-Resolution Depth Maps Based on TOF-Stereo Fusion
Vineet Gandhi, Jan Cech, Radu Horaud
The combination of range sensors with color cameras can be very useful for robot navigation, semantic perception, manipulation, and telepresence. Several methods of combining range…
No Cost Likelihood Manipulation at Test Time for Making Better Mistakes in Deep Networks
Shyamgopal Karthik, Ameya Prabhu, Puneet K. Dokania +1
There has been increasing interest in building deep hierarchy-aware classifiers that aim to quantify and reduce the severity of mistakes, and not just reduce the number of errors.…
ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction
Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti +3
We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network…
GAZED- Gaze-guided Cinematic Editing of Wide-Angle Monocular Video Recordings
K L Bhanu Moorthy, Moneish Kumar, Ramanathan Subramaniam +1
We present GAZED- eye GAZe-guided EDiting for videos captured by a solitary, static, wide-angle and high-resolution camera. Eye-gaze has been effectively employed in computational…