3 papers
cs.CV2026
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Shresth Grover, Priyank Pathak, Akash Kumar +1
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexpl…
cs.CV2026
How Do Inpainting Artifacts Propagate to Language?
Pratham Yashwante, Davit Abrahamyan, Shresth Grover +1
We study how visual artifacts introduced by diffusion-based inpainting affect language generation in vision-language models. We use a two-stage diagnostic setup in which masked ima…
cs.RO2025
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
Shresth Grover, Akshay Gopalkrishnan, Bo Ai +3
Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across di…