activity
20222024
most citedVision Transformer with Deformable Attention

20 citations · 31 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2024

Language Models Can Reduce Asymmetry in Information Markets

Nasim Rahaman, Martin Weiss, Manuel Wüthrich +4

This work addresses the buyer's inspection paradox for information markets. The paradox is that buyers need to access information to determine its value, while sellers need to limi…

cs.CV2024

Learning 3D object-centric representation through prediction

John Day, Tushar Arora, Jirui Liu +2

As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans…

cs.CL2023

The Role of Linguistic Priors in Measuring Compositional Generalization of Vision-Language Models

Chenwei Wu, Li Erran Li, Stefano Ermon +3

Compositionality is a common property in many modalities including natural languages and images, but the compositional generalization of multi-modal models is not well-understood.…

cs.CV202310 cited

DAT++: Spatially Dynamic Vision Transformer with Deformable Attention

Zhuofan Xia, Xuran Pan, Shiji Song +2

Transformers have shown superior performance on various vision tasks. Their large receptive field endows Transformer models with higher representation power than their CNN counterp…

cs.CV20231 cited

LiDAR-Based 3D Object Detection via Hybrid 2D Semantic Scene Generation

Haitao Yang, Zaiwei Zhang, Xiangru Huang +5

Bird's-Eye View (BEV) features are popular intermediate scene representations shared by the 3D backbone and the detector head in LiDAR-based object detectors. However, little resea…

cs.CV202220 cited

Vision Transformer with Deformable Attention

Zhuofan Xia, Xuran Pan, Shiji Song +2

Transformers have recently shown superior performances on various vision tasks. The large, sometimes even global, receptive field endows Transformer models with higher representati…