3 papers
cs.RO2026
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Dwip Dalal, Shivansh Patel, Chahit Jain +7
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. Howe…
cs.CV2025
City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
Dwip Dalal, Utkarsh Mishra, Narendra Ahuja +1
Leveraging multimodal large language models (MLLMs) to develop embodied agents offers significant promise for addressing complex real-world tasks. However, current evaluation bench…
cs.CV2025
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra +6
Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce At…