papers

Publications (44)

cs.LG2025

ControllableGPT: A Ground-Up Designed Controllable GPT for Molecule Optimization

Xuefeng Liu, Songhao Jiang, Bo Li +1

Large Language Models (LLMs) employ three popular training approaches: Masked Language Models (MLM), Causal Language Models (CLM), and Sequence-to-Sequence Models (seq2seq). Howeve…

q-bio.MN2023

Causal Discovery and Optimal Experimental Design for Genome-Scale Biological Network Recovery

Ashka Shah, Arvind Ramanathan, Valerie Hayot-Sasson +1

Causal discovery of genome-scale networks is important for identifying pathways from genes to observable traits - e.g. differences in cell function, disease, drug resistance and ot…

cs.LG2025

FragmentGPT: A Unified GPT Model for Fragment Growing, Linking, and Merging in Molecular Design

Xuefeng Liu, Songhao Jiang, Qinan Huang +5

Fragment-Based Drug Discovery (FBDD) is a popular approach in early drug development, but designing effective linkers to combine disconnected molecular fragments into chemically an…

cs.RO2026

AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots

Priyanka V. Setty, Arvind Ramanathan, Ian Foster +1

Self-driving laboratories increasingly rely on low-cost liquid handlers such as the Opentrons OT-2, which ship without the pressure-based aspiration monitoring of Hamilton or Tecan…

cs.AI2018

Precision Medicine as an Accelerator for Next Generation Cognitive Supercomputing

Edmon Begoli, Jim Brase, Bambi DeLaRosa +6

In the past several years, we have taken advantage of a number of opportunities to advance the intersection of next generation high-performance computing AI and big data technologi…

q-bio.QM2020

Ensemble Transfer Learning for the Prediction of Anti-Cancer Drug Response

Yitan Zhu, Thomas Brettin, Yvonne A. Evrard +6

Transfer learning has been shown to be effective in many applications in which training data for the target problem are limited but data for a related (source) problem are abundant…

cs.DC2021

Pandemic Drugs at Pandemic Speed: Infrastructure for Accelerating COVID-19 Drug Discovery with Hybrid Machine Learning- and Physics-based Simulations on High Performance Computers

Agastya P. Bhati, Shunzhou Wan, Dario Alfè +26

The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottle…

cs.LG2025

DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization

Xuefeng Liu, Songhao Jiang, Siyu Chen +4

Finetuning a Large Language Model (LLM) is crucial for generating results towards specific objectives. This research delves into the realm of drug optimization and introduce a nove…

q-bio.QM2021

Scaffold-Induced Molecular Graph (SIMG): Effective Graph Sampling Methods for High-Throughput Computational Drug Discovery

Austin Clyde, Ashka Shah, Max Zvyagin +2

Scaffold based drug discovery (SBDD) is a technique for drug discovery which pins chemical scaffolds as the framework of design. Scaffolds, or molecular frameworks, organize the de…

cs.LG2021

Learning from learning machines: a new generation of AI technology to meet the needs of science

Luca Pion-Tonachini, Kristofer Bouchard, Hector Garcia Martin +33

We outline emerging opportunities and challenges to enhance the utility of AI for scientific discovery. The distinct goals of AI for industry versus the goals of AI for science cre…

q-bio.BM2021

Protein-Ligand Docking Surrogate Models: A SARS-CoV-2 Benchmark for Deep Learning Accelerated Virtual Screening

Austin Clyde, Thomas Brettin, Alexander Partin +8

We propose a benchmark to study surrogate model accuracy for protein-ligand docking. We share a dataset consisting of 200 million 3D complex structures and 2D structure scores acro…

cs.RO2026

PRISM: Protocol Refinement through Intelligent Simulation Modeling

Brian Hsu, Priyanka V Setty, Rory M Butler +7

Automating experimental protocol design and execution remains as a fundamental bottleneck in realizing self-driving laboratories. We introduce PRISM (Protocol Refinement through In…

cs.CV2025

Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis

Robinson Umeike, Neil Getty, Fangfang Xia +1

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researchers accelerate scientific disco…

cs.DC2020

Scalable HPC and AI Infrastructure for COVID-19 Therapeutics

Hyungro Lee, Andre Merzky, Li Tan +15

COVID-19 has claimed more 1 million lives and resulted in over 40 million infections. There is an urgent need to identify drugs that can inhibit SARS-CoV-2. In response, the DOE re…

cs.LG2019

Scalable Reinforcement-Learning-Based Neural Architecture Search for Cancer Deep Learning Research

Prasanna Balaprakash, Romain Egele, Misha Salim +5

Cancer is a complex disease, the understanding and treatment of which are being aided through increases in the volume of collected data and in the scale of deployed computing power…

cs.IR2025

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace +21

The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval Augm…

q-bio.BM2020

Targeting SARS-CoV-2 with AI- and HPC-enabled Lead Generation: A First Data Release

Yadu Babuji, Ben Blaiszik, Tom Brettin +15

Researchers across the globe are seeking to rapidly repurpose existing drugs or discover new drugs to counter the the novel coronavirus disease (COVID-19) caused by severe acute re…

cs.LG2026

Monte Carlo Tree Diffusion with Multiple Experts for Protein Design

Xuefeng Liu, Mingxuan Cao, Songhao Jiang +6

The goal of protein design is to generate amino acid sequences that fold into functional structures with desired properties. Prior methods combining autoregressive language models…

eess.IV2020

Deep Medical Image Analysis with Representation Learning and Neuromorphic Computing

Neil Getty, Thomas Brettin, Dong Jin +2

We explore three representative lines of research and demonstrate the utility of our methods on a classification benchmark of brain cancer MRI data. First, we present a capsule net…

cs.LG2025

Causal Discovery over High-Dimensional Structured Hypothesis Spaces with Causal Graph Partitioning

Ashka Shah, Adela DePavia, Nathaniel Hudson +2

The aim in many sciences is to understand the mechanisms that underlie the observed distribution of variables, starting from a set of initial hypotheses. Causal discovery allows us…

q-bio.GN2026

Causal Discovery of Radiation Response Mechanisms in Human Cells

Ashka Shah, Rick Stevens

The paper applies causal discovery methods to RNA‑sequencing data to infer directed gene regulatory networks that explain how human cells respond to different radiation dose rates.

#causal discovery#radiation response#gene expression#network inference
cs.AI2023

DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang +89

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scient…

q-bio.QM2021

A cross-study analysis of drug response prediction in cancer cell lines

Fangfang Xia, Jonathan Allen, Prasanna Balaprakash +21

To enable personalized cancer treatment, machine learning models have been developed to predict drug response as a function of tumor and drug features. However, most algorithm deve…

cs.LG2026

Active Advantage-Aligned Online Reinforcement Learning with Offline Data

Xuefeng Liu, Hung T. C. Le, Siyu Chen +4

Online reinforcement learning (RL) enhances policies through direct interactions with the environment, but faces challenges related to sample efficiency. In contrast, offline RL le…

cs.LG2024

Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision

Nathaniel Hudson, J. Gregory Pauloski, Matt Baughman +13

Deep learning methods are transforming research, enabling new techniques, and ultimately leading to new discoveries. As the demand for more capable AI models continues to grow, we…

cs.AI2025

Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling

Peng Ding, Rick Stevens

The proliferation of tool-augmented Large Language Models (LLMs) has created a fragmented ecosystem where developers must navigate multiple protocols, manual schema definitions, an…

cs.IR2025

AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine

Carlo Siebenschuh, Kyle Hippe, Ozan Gokdemir +10

Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive…

cs.SE2026

Stdlib or Third-Party? Empirical Performance and Correctness of LLM-Assisted Zero-Dependency Python Libraries

Peng Ding, Rick Stevens

Third-party Python libraries introduce dependency management overhead, supply chain risk, and deployment friction in constrained environments. A natural question is how much of thi…

cs.DC2020

IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads

Aymen Al Saadi, Dario Alfe, Yadu Babuji +33

The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive…

cs.SE2026

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

Peng Ding, Rick Stevens

Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MCP, OpenAPI, LangChain) is integrated f…

q-bio.QM2020

Learning Curves for Drug Response Prediction in Cancer Cell Lines

Alexander Partin, Thomas Brettin, Yvonne A. Evrard +9

Motivated by the size of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As…

stat.ML2016

Machine Learning for Antimicrobial Resistance

John W. Santerre, James J. Davis, Fangfang Xia +1

Biological datasets amenable to applied machine learning are more available today than ever before, yet they lack adequate representation in the Data-for-Good community. Here we pr…

q-bio.BM2025

ScaffoldGPT: A Scaffold-based GPT Model for Drug Optimization

Xuefeng Liu, Songhao Jiang, Ian Foster +2

Drug optimization has become increasingly crucial in light of fast-mutating virus strains and drug-resistant cancer cells. Nevertheless, it remains challenging as it necessitates r…

cs.LG2025

Bidirectional Hierarchical Protein Multi-Modal Representation Learning

Xuefeng Liu, Songhao Jiang, Chih-chan Tien +2

Protein representation learning is critical for numerous biological tasks. Recently, large transformer-based protein language models (pLMs) pretrained on large scale protein sequen…

q-bio.QM2020

Regression Enrichment Surfaces: a Simple Analysis Technique for Virtual Drug Screening Models

Austin Clyde, Xiaotian Duan, Rick Stevens

We present a new method for understanding the performance of a model in virtual drug screening tasks. While most virtual screening problems present as a mix between ranking and cla…

cs.LG2020

A Systematic Approach to Featurization for Cancer Drug Sensitivity Predictions with Deep Learning

Austin Clyde, Tom Brettin, Alexander Partin +6

By combining various cancer cell line (CCL) drug screening panels, the size of the data has grown significantly to begin understanding how advances in deep learning can advance dru…

cs.DC2025

Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery

William E. Allcock, Benjamin S. Allen, James Anchell +106

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technolo…

cs.AI2025

EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Franck Cappello, Sandeep Madireddy, Robert Underwood +23

Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that req…

cs.LG2023

WordScape: a Pipeline to extract multilingual, visually rich Documents with Layout Annotations from Web Crawl Data

Maurice Weber, Carlo Siebenschuh, Rory Butler +8

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. R…

cs.LG2021

Neko: a Library for Exploring Neuromorphic Learning Rules

Zixuan Zhao, Nathan Wycoff, Neil Getty +2

The field of neuromorphic computing is in a period of active exploration. While many tools have been developed to simulate neuronal dynamics or convert deep networks to spiking mod…

cs.LG2021

Scaffold Embeddings: Learning the Structure Spanned by Chemical Fragments, Scaffolds and Compounds

Austin Clyde, Arvind Ramanathan, Rick Stevens

Molecules have seemed like a natural fit to deep learning's tendency to handle a complex structure through representation learning, given enough data. However, this often continuou…

cs.LG2026

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

Xuefeng Liu, Mingxuan Cao, Qinan Huang +3

Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open cha…

cs.RO2023

Towards a Modular Architecture for Science Factories

Rafael Vescovi, Tobias Ginsburg, Kyle Hippe +14

Advances in robotic automation, high-performance computing (HPC), and artificial intelligence (AI) encourage us to conceive of science factories: large, general-purpose computation…

q-bio.QM2026

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor +2

Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. How…