activity
20242026
collaborators

8 papers

cs.LG2026

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models

Yangyue Wang, Harshvardhan Sikka, Yash Mathur +3

GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming…

cs.CR2026

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

Kun Wang, Cheng Qian, Miao Yu +6

Multimodal Large Language Models (MLLMs) have achieved remarkable success in cross-modal understanding and generation, yet their deployment is threatened by critical safety vulnera…

cs.CV2026

Online Reasoning Video Object Segmentation

Jinyuan Liu, Yang Wang, Zeyu Zhao +3

Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existi…

cs.RO2026

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

Xiaolei Lang, Yang Wang, Yukun Zhou +10

Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-world tasks. However, scaling th…

cs.LG2025

Benchmarking the Generality of Vision-Language-Action Models

Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka +6

Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices r…

cs.LG2025

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

Pranav Guruprasad, Yangyue Wang, Sudipta Chowdhury +2

Recent innovations in multimodal action models represent a promising direction for developing general-purpose agentic systems, combining visual understanding, language comprehensio…