activity
20202022
most citedWhat Makes Multi-modal Learning Better than Single (Provably)

43 citations · 73 across the 7 of their papers we have counts for

collaborators

11 papers

cs.CV2022

Learning Visual Styles from Audio-Visual Associations

Tingle Li, Yichen Liu, Andrew Owens +1

From the patter of rain to the crunch of snow, the sounds we hear often convey the visual textures that appear within a scene. In this paper, we present a method for learning visua…

cs.CV2022

MUTR3D: A Multi-camera Tracking Framework via 3D-to-2D Queries

Tianyuan Zhang, Xuanyao Chen, Yue Wang +2

Accurate and consistent 3D tracking from multiple cameras is a key component in a vision-based autonomous driving system. It involves modeling 3D dynamic objects in complex scenes…

cs.CV2022

Training-Free Robust Multimodal Learning via Sample-Wise Jacobian Regularization

Zhengqi Gao, Sucheng Ren, Zihui Xue +2

Multimodal fusion emerges as an appealing technique to improve model performances on many tasks. Nevertheless, the robustness of such fusion methods is rarely involved in the prese…

cs.CV20212 cited

Co-advise: Cross Inductive Bias Distillation

Sucheng Ren, Zhengqi Gao, Tianyu Hua +4

Transformers recently are adapted from the community of natural language processing as a promising substitute of convolution-based neural networks for visual learning tasks. Howeve…

cs.LG202125 cited

Improving Multi-Modal Learning with Uni-Modal Teachers

Chenzhuang Du, Tingle Li, Yichen Liu +4

Learning multi-modal representations is an essential step towards real-world robotic applications, and various multi-modal fusion models have been developed for this purpose. Howev…

cs.LG202143 cited

What Makes Multi-modal Learning Better than Single (Provably)

Yu Huang, Chenzhuang Du, Zihui Xue +3

The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is…