From the 1 of 9 linked papers with an AI index.
9 papers
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7
LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…
T-Rex: Tactile-Reactive Dexterous Manipulation
Dantong Niu, Zhuoyang Liu, Zekai Wang +31
The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…
Contrastive Action-Image Pre-training for Visuomotor Control
Yuvan Sharma, Dantong Niu, Anirudh Pai +16
Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarci…
Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2
Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including mathematical reasoning and text-…
Learning to Grasp Anything by Playing with Random Toys
Dantong Niu, Yuvan Sharma, Baifeng Shi +11
Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop genera…
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +3
World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face prac…