works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7

LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…

cs.RO2026

T-Rex: Tactile-Reactive Dexterous Manipulation

Dantong Niu, Zhuoyang Liu, Zekai Wang +31

The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…

cs.RO2026

Contrastive Action-Image Pre-training for Visuomotor Control

Yuvan Sharma, Dantong Niu, Anirudh Pai +16

Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarci…

cs.CV2026

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2

Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including mathematical reasoning and text-…

cs.RO2026

Learning to Grasp Anything by Playing with Random Toys

Dantong Niu, Yuvan Sharma, Baifeng Shi +11

Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop genera…

cs.AI2025

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +3

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face prac…