works on

From the 1 of 36 linked papers with an AI index.

activity
20242026
most citedRoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Li Siyao, Jiawei Gu +16

The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…

cs.CV2026

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna +1

Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains…

cs.CV2026

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

Tanush Yadav, Mohammadreza Salehi, Jae Sung Park +6

Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessential task for video understa…

cs.CV2026

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

Tanmay Gupta, Piper Wolters, Zixian Ma +13

Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, t…

cs.CV2025

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

Henry Herzog, Favyen Bastani, Yawen Zhang +23

Earth observation data presents a unique challenge: it is spatial like images, sequential like video or text, and highly multimodal. We present OlmoEarth: a multimodal, spatio-temp…

cs.CV2025

One Diffusion to Generate Them All

Duong H. Le, Tuan Pham, Sangho Lee +5

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables condit…