activity
20202025
most citedWhite-Box Transformers via Sparse Rate Reduction

23 citations · 46 across the 14 of their papers we have counts for

collaborators

15 papers

cs.CV2025

Language-Image Alignment with Fixed Text Encoders

Jingfeng Yang, Ziyang Wu, Yue Zhao +1

Currently, the most dominant approach to establishing language-image alignment is to pre-train text and image encoders jointly through contrastive learning, such as CLIP and its va…

cs.CV2025

Simplifying DINO via Coding Rate Regularization

Ziyang Wu, Jingyuan Zhang, Druv Pai +5

DINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-t…

cs.LG2024★ 7 cited

Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

Ziyang Wu, Tianjiao Ding, Yifu Lu +6

The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…

cs.LG2024

Spatial-Temporal Mixture-of-Graph-Experts for Multi-Type Crime Prediction

Ziyang Wu, Fan Liu, Jindong Han +2

As various types of crime continue to threaten public safety and economic development, predicting the occurrence of multiple types of crimes becomes increasingly vital for effectiv…

cs.SD2024★ 2 cited

A Survey of Foundation Models for Music Understanding

Wenjun Li, Ying Cai, Ziyang Wu +13

Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance…

cs.LG2024

Masked Completion via Structured Diffusion with White-Box Transformers

Druv Pai, Ziyang Wu, Sam Buchanan +2

Modern learning frameworks often train deep neural networks with massive amounts of unlabeled data to learn representations by solving simple pretext tasks, then use the representa…