activity
20182026
most citedLearning Object Detection from Captions via Textual Scene Attributes

6 citations · 8 across the 17 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.RO2026

ContextFlow: In-Context Flow Matching for Robot Manipulation

Jian Ding, Xianjie Dai, Roei Herzig +5

Although highly effective in vision and language domains, applying in-context learning to robotics remains challenging. Existing autoregressive in-context imitation methods discret…

cs.AI2026

Discriminative World Models for Web Agents

Kelvin Li, Dhruv Pendharkar, Anish Pahilajani +6

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Proc…

cs.RO2026

T-Rex: Tactile-Reactive Dexterous Manipulation

Dantong Niu, Zhuoyang Liu, Zekai Wang +31

The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…

cs.RO2026

Playful Agentic Robot Learning

Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reu…

cs.RO2026

Contrastive Action-Image Pre-training for Visuomotor Control

Yuvan Sharma, Dantong Niu, Anirudh Pai +16

Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarci…

cs.CV2026

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

Jovana Kondic, Pengyuan Li, Dhiraj Joshi +24

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language…