collaborators

6 papers

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…

cs.LG2026

Same Graph, Different Likelihoods: Calibration of Autoregressive Graph Generators via Permutation-Equivalent Encodings

Laurits Fredsgaard, Aaron Thomas, Michael Riis Andersen +2

Autoregressive graph generators define likelihoods via a sequential construction process, but these likelihoods are only meaningful if they are consistent across all linearizations…

quant-ph2025

QCA-MolGAN: Quantum Circuit Associative Molecular GAN with Multi-Agent Reinforcement Learning

Aaron Mark Thomas, Yu-Cheng Chen, Hubert Okadome Valencia +2

Navigating the vast chemical space of molecular structures to design novel drug molecules with desired target properties remains a central challenge in drug discovery. Recent advan…

cs.CL2025

Refining Salience-Aware Sparse Fine-Tuning Strategies for Language Models

Xinxin Liu, Aaron Thomas, Cheng Zhang +3

Parameter-Efficient Fine-Tuning (PEFT) has gained prominence through low-rank adaptation methods like LoRA. In this paper, we focus on sparsity-based PEFT (SPEFT), which introduces…

quant-ph2025

VAE-QWGAN: Addressing Mode Collapse in Quantum GANs via Autoencoding Priors

Aaron Mark Thomas, Harry Youel, Sharu Theresa Jose

Recent proposals for quantum generative adversarial networks (GANs) suffer from the issue of mode collapse, analogous to classical GANs, wherein the distribution learnt by the GAN…

quant-ph2025

On the Generalization of Adversarially Trained Quantum Classifiers

Petros Georgiou, Aaron Mark Thomas, Sharu Theresa Jose +1

Quantum classifiers are vulnerable to adversarial attacks that manipulate their input classical or quantum data. A promising countermeasure is adversarial training, where quantum c…