most citedMake A Long Image Short: Adaptive Token Length for Vision Transformers

2 citations · 2 across the 2 of their papers we have counts for

collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2024

MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

Jinming Li, Yichen Zhu, Zhiyuan Xu +7

It is fundamentally challenging for robots to serve as useful assistants in human environments because this requires addressing a spectrum of sub-problems across robotics, includin…

cs.RO2024

Language-Conditioned Robotic Manipulation with Fast and Slow Thinking

Minjie Zhu, Yichen Zhu, Jinming Li +8

The language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple pick-and-place to tasks requiring intent recogniti…

cs.RO2024

Object-Centric Instruction Augmentation for Robotic Manipulation

Junjie Wen, Yichen Zhu, Minjie Zhu +8

Humans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as \enquote{pick and place}, understand…

cs.RO2024

Visual Robotic Manipulation with Depth-Aware Pretraining

Wanying Wang, Jinming Li, Yichen Zhu +7

Recent work on visual representation learning has shown to be efficient for robotic manipulation tasks. However, most existing works pretrained the visual backbone solely on 2D ima…

cs.RO20231 cited

CMG-Net: An End-to-End Contact-Based Multi-Finger Dexterous Grasping Network

Mingze Wei, Yaomin Huang, Zhiyuan Xu +7

In this paper, we propose a novel representation for grasping using contacts between multi-finger robotic hands and objects to be manipulated. This representation significantly red…