11 papers
Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
David R. Wessels, Farhad Ramezanghorbani, David W. Romero +9
Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while re…
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +571
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…
Probabilistic Contrastive Pretraining for Multi-task ADME Property Prediction
Yifan Xue, Srimukh Prasad Veccham, Saee Paliwal +2
Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy…
Fold-CP: A Context Parallelism Framework for Biomolecular Modeling
Dejun Lin, Simon Chu, Vishanth Iyer +35
Understanding cellular machinery requires atomic-scale reconstruction of large biomolecular assemblies. However, predicting the structures of these systems has been constrained by…
Exploring Synthesizable Chemical Space with Iterative Pathway Refinements
Seul Lee, Karsten Kreis, Srimukh Prasad Veccham +5
A well-known pitfall of molecular generative models is that they are not guaranteed to generate synthesizable molecules. Existing solutions for this problem often struggle to effec…
Multitask finetuning and acceleration of chemical pretrained models for small molecule drug property prediction
Matthew Adrian, Yunsie Chung, Kevin Boyd +3
Chemical pretrained models, sometimes referred to as foundation models, are receiving considerable interest for drug discovery applications. The general chemical knowledge extracte…