activity
20142024
most citedRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

273 citations · 364 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Depth Map Denoising Network and Lightweight Fusion Network for Enhanced 3D Face Recognition

Ruizhuo Xu, Ke Wang, Chao Deng +5

With the increasing availability of consumer depth sensors, 3D face recognition (FR) has attracted more and more attention. However, the data acquired by these sensors are often co…

cs.CV202326 cited

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…

cs.CV202339 cited

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…

cs.CV20236 cited

ScribbleSeg: Scribble-based Interactive Image Segmentation

Xi Chen, Yau Shing Jonathan Cheung, Ser-Nam Lim +1

Interactive segmentation enables users to extract masks by providing simple annotations to indicate the target, such as boxes, clicks, or scribbles. Among these interaction formats…

cs.CV20235 cited

Knowledge-augmented Few-shot Visual Relation Detection

Tianyu Yu, Yangning Li, Jiaoyan Chen +8

Visual Relation Detection (VRD) aims to detect relationships between objects for image understanding. Most existing VRD methods rely on thousands of training samples of each relati…