collaborators

7 papers

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…

cs.LG2026

Probabilistic Contrastive Pretraining for Multi-task ADME Property Prediction

Yifan Xue, Srimukh Prasad Veccham, Saee Paliwal +2

Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy…

cs.LG2026

Exploring Synthesizable Chemical Space with Iterative Pathway Refinements

Seul Lee, Karsten Kreis, Srimukh Prasad Veccham +5

A well-known pitfall of molecular generative models is that they are not guaranteed to generate synthesizable molecules. Existing solutions for this problem often struggle to effec…

cs.LG2025

Multitask finetuning and acceleration of chemical pretrained models for small molecule drug property prediction

Matthew Adrian, Yunsie Chung, Kevin Boyd +3

Chemical pretrained models, sometimes referred to as foundation models, are receiving considerable interest for drug discovery applications. The general chemical knowledge extracte…

cs.LG2025

Want to train KANS at scale? Now UKAN!

Alireza Moradzadeh, Srimukh Prasad Veccham, Lukasz Wawrzyniak +2

Kolmogorov-Arnold Networks (KANs) have recently emerged as a powerful alternative to traditional multilayer perceptrons. However, their reliance on predefined, bounded grids restri…

cs.LG2025

BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery

Peter St. John, Dejun Lin, Polina Binder +89

Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increas…