8 papers
Masked Language Flow Models
Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2
Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in th…
Variance-Tilted Diffusion Models for Diverse Sampling
Iskander Azangulov, Leo Zhang, Kianoosh Ashouritaklimi
Diffusion models are typically sampled independently, even when the downstream objective is to obtain a diverse set of candidates. We introduce a variance-weighted batch distributi…
Towards a Frugal Photosynthesis Sensing Toolkit for Data-Driven Plant Science Education and Exploration
Qitong Li, Raj Nileshbhai Dave, Rhema Amanda Phiri +7
Rapid environmental change and advances in data-driven analysis highlight the need not only to use computational tools, but also to foster understanding of the natural world and in…
SigmaDock: Untwisting Molecular Docking With Fragment-Based SE(3) Diffusion
Alvaro Prat, Leo Zhang, Charlotte M. Deane +2
Determining the binding pose of a ligand to a protein, known as molecular docking, is a fundamental task in drug discovery. Generative approaches promise faster, improved, and more…
Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLM
YuanLab. ai, :, Shawn Wu +25
We introduce Yuan3.0 Ultra, an open-source Mixture-of-Experts (MoE) large language model featuring 68.8B activated parameters and 1010B total parameters, specially designed to enha…
Orthogonal Self-Attention
Leo Zhang, James Martens
Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, re…