most citedEfficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications

9 citations · 15 across the 3 of their papers we have counts for

collaborators

7 papers

cs.RO2025

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Zhi Hou, Tianyi Zhang, Yuwen Xiong +8

While recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact ac…

cs.CV2025

LangBridge: Interpreting Image as a Combination of Language Embeddings

Jiaqi Liao, Yuwei Niu, Fanqing Meng +9

Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various complex vision-language tasks. F…

cs.CV2024

HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving

Zehuan Wu, Jingcheng Ni, Xiaodong Wang +5

Generative models have significantly improved the generation and prediction quality on either camera images or LiDAR point clouds for autonomous driving. However, a real-world auto…

cs.CV2024

big.LITTLE Vision Transformer for Efficient Visual Recognition

He Guo, Yulong Wang, Zixuan Ye +2

In this paper, we introduce the big.LITTLE Vision Transformer, an innovative architecture aimed at achieving efficient visual recognition. This dual-transformer system is composed…

cs.RO2024

Diffusion Transformer Policy

Zhi Hou, Tianyi Zhang, Yuwen Xiong +6

Recent large vision-language-action models pretrained on diverse robot datasets have demonstrated the potential for generalizing to new environments with a few in-domain data. Howe…

cs.CV20249 cited

Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications

Yuwen Xiong, Zhiqi Li, Yuntao Chen +10

We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of i…