43 citations · 73 across the 7 of their papers we have counts for
11 papers
Learning Visual Styles from Audio-Visual Associations
Tingle Li, Yichen Liu, Andrew Owens +1
From the patter of rain to the crunch of snow, the sounds we hear often convey the visual textures that appear within a scene. In this paper, we present a method for learning visua…
MUTR3D: A Multi-camera Tracking Framework via 3D-to-2D Queries
Tianyuan Zhang, Xuanyao Chen, Yue Wang +2
Accurate and consistent 3D tracking from multiple cameras is a key component in a vision-based autonomous driving system. It involves modeling 3D dynamic objects in complex scenes…
Training-Free Robust Multimodal Learning via Sample-Wise Jacobian Regularization
Zhengqi Gao, Sucheng Ren, Zihui Xue +2
Multimodal fusion emerges as an appealing technique to improve model performances on many tasks. Nevertheless, the robustness of such fusion methods is rarely involved in the prese…
Co-advise: Cross Inductive Bias Distillation
Sucheng Ren, Zhengqi Gao, Tianyu Hua +4
Transformers recently are adapted from the community of natural language processing as a promising substitute of convolution-based neural networks for visual learning tasks. Howeve…
Improving Multi-Modal Learning with Uni-Modal Teachers
Chenzhuang Du, Tingle Li, Yichen Liu +4
Learning multi-modal representations is an essential step towards real-world robotic applications, and various multi-modal fusion models have been developed for this purpose. Howev…
What Makes Multi-modal Learning Better than Single (Provably)
Yu Huang, Chenzhuang Du, Zihui Xue +3
The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is…