From the 1 of 25 linked papers with an AI index.
1 citations · 1 across the 2 of their papers we have counts for
25 papers
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Echoverse introduces a pipeline that compiles specifications into deep, stateful synthetic applications for training computer-use agents, using a co‑evolution loop that repairs env…
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of s…
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Shresth Grover, Priyank Pathak, Akash Kumar +1
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexpl…
Fara-1.5: Scalable Learning Environments for Computer Use Agents
Ahmed Awadallah, Sahil Gupta, Yash Lara +12
Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires two key ingredients: environment…
ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
Amirhossein Abaskohi, Yuhang He, Peter West +3
Computer-use agents (CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual tokens. As interaction trajector…
What MLLMs Learn about When they Learn about Multimodal Reasoning
Jiwan Chung, Neel Joshi, Pratyusha Sharma +2
Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark…