collaborators

8 papers

cs.CL2026

Masked Language Flow Models

Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2

Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in th…

stat.ML2026

Variance-Tilted Diffusion Models for Diverse Sampling

Iskander Azangulov, Leo Zhang, Kianoosh Ashouritaklimi

Diffusion models are typically sampled independently, even when the downstream objective is to obtain a diverse set of candidates. We introduce a variance-weighted batch distributi…

cs.HC2026

Towards a Frugal Photosynthesis Sensing Toolkit for Data-Driven Plant Science Education and Exploration

Qitong Li, Raj Nileshbhai Dave, Rhema Amanda Phiri +7

Rapid environmental change and advances in data-driven analysis highlight the need not only to use computational tools, but also to foster understanding of the natural world and in…

cs.LG2026

SigmaDock: Untwisting Molecular Docking With Fragment-Based SE(3) Diffusion

Alvaro Prat, Leo Zhang, Charlotte M. Deane +2

Determining the binding pose of a ligand to a protein, known as molecular docking, is a fundamental task in drug discovery. Generative approaches promise faster, improved, and more…

cs.LG2026

Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLM

YuanLab. ai, :, Shawn Wu +25

We introduce Yuan3.0 Ultra, an open-source Mixture-of-Experts (MoE) large language model featuring 68.8B activated parameters and 1010B total parameters, specially designed to enha…

cs.LG2026

Orthogonal Self-Attention

Leo Zhang, James Martens

Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, re…