6 papers
Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
David R. Wessels, Farhad Ramezanghorbani, David W. Romero +9
Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while re…
PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding
Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts +3
Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules.…
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +571
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…
Flash Invariant Point Attention
Andrew Liu, Axel Elaldi, Nicholas T Franklin +4
Invariant Point Attention (IPA) is a key algorithm for geometry-aware modeling in structural biology, central to many protein and RNA models. However, its quadratic complexity limi…
Bio2Token: All-atom tokenization of any biomolecular structure with Mamba
Andrew Liu, Axel Elaldi, Nathan Russell +1
Efficient encoding and representation of large 3D molecular structures with high fidelity is critical for biomolecular design applications. Despite this, many representation learni…
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization
Ryan Park, Darren J. Hsu, C. Brian Roland +5
Inverse folding models play an important role in structure-based design by predicting amino acid sequences that fold into desired reference structures. Models like ProteinMPNN, a m…