activity
20162022
most citedThe Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

143 citations · 576 across the 31 of their papers we have counts for

collaborators

49 papers

cs.CV2022

AutoGPart: Intermediate Supervision Search for Generalizable 3D Part Segmentation

Xueyi Liu, Xiaomeng Xu, Anyi Rao +2

Training a generalizable 3D part segmentation network is quite challenging but of great importance in real-world applications. To tackle this problem, some works design task-specif…

cs.CV20216 cited

When Does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?

Lijie Fan, Sijia Liu, Pin-Yu Chen +2

Contrastive learning (CL) can learn generalizable feature representations and achieve the state-of-the-art performance of downstream tasks by finetuning a linear classifier on top…

cs.CV20211 cited

TSM: Temporal Shift Module for Efficient and Scalable Video Understanding on Edge Device

Ji Lin, Chuang Gan, Kuan Wang +1

The explosive growth in video streaming requires video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture te…

cs.LG2021

Certifiably Robust Interpretation via Renyi Differential Privacy

Ao Liu, Xiaoyu Chen, Sijia Liu +2

Motivated by the recent discovery that the interpretation maps of CNNs could easily be manipulated by adversarial attacks against network interpretability, we study the problem of…

eess.AS202112 cited

Global Rhythm Style Transfer Without Text Transcriptions

Kaizhi Qian, Yang Zhang, Shiyu Chang +4

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody…

cs.CV20215 cited

Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning

Shaobo Min, Qi Dai, Hongtao Xie +3

Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual an…