3 papers
cs.CV2026
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
Niyati Rawal, Sushant Ravva, Shah Alam Abir +5
Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spati…
cs.RO2024
UNMuTe: Unifying Navigation and Multimodal Dialogue-like Text Generation
Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi +1
Smart autonomous agents are becoming increasingly important in various real-life applications, including robotics and autonomous vehicles. One crucial skill that these agents must…
cs.CV2024
AIGeN: An Adversarial Approach for Instruction Generation in VLN
Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi +1
In the last few years, the research interest in Vision-and-Language Navigation (VLN) has grown significantly. VLN is a challenging task that involves an agent following human instr…