273 citations · 364 across the 16 of their papers we have counts for
5 papers · 1 filter
Depth Map Denoising Network and Lightweight Fusion Network for Enhanced 3D Face Recognition
Ruizhuo Xu, Ke Wang, Chao Deng +5
With the increasing availability of consumer depth sensors, 3D face recognition (FR) has attracted more and more attention. However, the data acquired by these sensors are often co…
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Xi Chen, Xiao Wang, Lucas Beyer +16
This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Xi Chen, Josip Djolonga, Piotr Padlewski +40
We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…
ScribbleSeg: Scribble-based Interactive Image Segmentation
Xi Chen, Yau Shing Jonathan Cheung, Ser-Nam Lim +1
Interactive segmentation enables users to extract masks by providing simple annotations to indicate the target, such as boxes, clicks, or scribbles. Among these interaction formats…
Knowledge-augmented Few-shot Visual Relation Detection
Tianyu Yu, Yangning Li, Jiaoyan Chen +8
Visual Relation Detection (VRD) aims to detect relationships between objects for image understanding. Most existing VRD methods rely on thousands of training samples of each relati…