9 papers
Language-Image Alignment with Fixed Text Encoders
Jingfeng Yang, Ziyang Wu, Yue Zhao +1
Currently, the most dominant approach to establishing language-image alignment is to pre-train text and image encoders jointly through contrastive learning, such as CLIP and its va…
Simplifying DINO via Coding Rate Regularization
Ziyang Wu, Jingyuan Zhang, Druv Pai +5
DINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-t…
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…
LLoCO: Learning Long Contexts Offline
Sijun Tan, Xiuyu Li, Shishir Patil +5
Processing long contexts remains a challenge for large language models (LLMs) due to the quadratic computational and memory overhead of the self-attention mechanism and the substan…
Spatial-Temporal Mixture-of-Graph-Experts for Multi-Type Crime Prediction
Ziyang Wu, Fan Liu, Jindong Han +2
As various types of crime continue to threaten public safety and economic development, predicting the occurrence of multiple types of crimes becomes increasingly vital for effectiv…
A Survey of Foundation Models for Music Understanding
Wenjun Li, Ying Cai, Ziyang Wu +13
Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance…