most citedWhite-Box Transformers via Sparse Rate Reduction

23 citations · 35 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV20247 cited

Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Shengbang Tong, Zhuang Liu, Yuexiang Zhai +3

Is vision good enough for language? Recent advancements in multimodal models primarily stem from the powerful reasoning abilities of large language models (LLMs). However, the visu…

cs.CV20233 cited

Emergence of Segmentation with Minimalistic White-Box Transformers

Yaodong Yu, Tianzhe Chu, Shengbang Tong +4

Transformer-like models for vision tasks have recently proven effective for a wide range of downstream applications such as segmentation and detection. Previous works have shown th…

cs.LG202323 cited

White-Box Transformers via Sparse Rate Reduction

Yaodong Yu, Sam Buchanan, Druv Pai +5

In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dime…

cs.CV202311 cited

EMP-SSL: Towards Self-Supervised Learning in One Training Epoch

Shengbang Tong, Yubei Chen, Yi Ma +1

Recently, self-supervised learning (SSL) has achieved tremendous success in learning image representation. Despite the empirical success, most self-supervised learning methods are…

cs.CV20231 cited

Closed-Loop Transcription via Convolutional Sparse Coding

Xili Dai, Ke Chen, Shengbang Tong +9

Autoencoding has achieved great empirical success as a framework for learning generative models for natural images. Autoencoders often use generic deep networks as the encoder or d…