activity
20242026
collaborators

6 papers

cs.LG2026

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models

Yangyue Wang, Harshvardhan Sikka, Yash Mathur +3

GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming…

cs.LG2025

Benchmarking the Generality of Vision-Language-Action Models

Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka +6

Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices r…

cs.LG2025

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

Pranav Guruprasad, Yangyue Wang, Sudipta Chowdhury +2

Recent innovations in multimodal action models represent a promising direction for developing general-purpose agentic systems, combining visual understanding, language comprehensio…

cs.CV2025

Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments

Pranav Guruprasad, Yangyue Wang, Sudipta Chowdhury +2

Vision-language-action (VLA) models represent an important step toward general-purpose robotic systems by integrating visual perception, language understanding, and action executio…

cs.RO2024

Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

Pranav Guruprasad, Harshvardhan Sikka, Jaewoo Song +2

Vision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, langu…

cs.CL2024

KULCQ: An Unsupervised Keyword-based Utterance Level Clustering Quality Metric

Pranav Guruprasad, Negar Mokhberian, Nikhil Varghese +2

Intent discovery is crucial for both building new conversational agents and improving existing ones. While several approaches have been proposed for intent discovery, most rely on…