2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.LG2026
Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs
Wentao Lu, Jesse Clark, Tianyu Zhu
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods dif…
cs.IR2024★ 2 cited
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
Tianyu Zhu, Myong Chol Jung, Jesse Clark
Contrastive learning has gained widespread adoption for retrieval tasks due to its minimal requirement for manual annotations. However, popular training frameworks typically learn…