works on

From the 1 of 25 linked papers with an AI index.

most citedOn Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

1 citations · 1 across the 2 of their papers we have counts for

collaborators

25 papers

cs.AI2026

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Yash Pandya, Sahil Gupta, Sarthak Harne +10

Echoverse introduces a pipeline that compiles specifications into deep, stateful synthetic applications for training computer-use agents, using a co‑evolution loop that repairs env…

cs.CV2026

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat

This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of s…

cs.CV2026

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

Shresth Grover, Priyank Pathak, Akash Kumar +1

Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexpl…

cs.AI2026

Fara-1.5: Scalable Learning Environments for Computer Use Agents

Ahmed Awadallah, Sahil Gupta, Yash Lara +12

Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires two key ingredients: environment…

cs.CL2026

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

Amirhossein Abaskohi, Yuhang He, Peter West +3

Computer-use agents (CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual tokens. As interaction trajector…

cs.CL2026

What MLLMs Learn about When they Learn about Multimodal Reasoning

Jiwan Chung, Neel Joshi, Pratyusha Sharma +2

Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark…