collaborators

6 papers

q-bio.QM2026

Scalable Agentic Reasoning for Designing Biologics Targeting Intrinsically Disordered Proteins

Matthew Sinclair, Moeen Meigooni, Archit Vasan +14

Intrinsically disordered proteins (IDPs) represent crucial therapeutic targets due to their significant role in disease -- approximately 80\% of cancer-related proteins contain lon…

cs.AI2026

BioAlchemy: Distilling Biological Literature into Reasoning-Ready Reinforcement Learning Training Data

Brian Hsu, Ozan Gökdemir, Carlo Siebenschuh +7

Despite the large corpus of biology training text, the impact of reasoning models on biological research generally lags behind math and coding. In this work, we show that biology q…

cs.LG2026

LSHBloom: Memory-efficient, Extreme-scale Document Deduplication

Arham Khan, Robert Underwood, Carlo Siebenschuh +7

Contemporary large language model (LLM) training pipelines require the assembly of internet-scale databases full of text data from a variety of sources (e.g., web, academic, and pu…

cs.IR2025

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace +21

The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval Augm…

cs.IR2025

AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine

Carlo Siebenschuh, Kyle Hippe, Ozan Gokdemir +10

Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive…

cs.DC2025

Connecting Large Language Model Agent to High Performance Computing Resource

Heng Ma, Alexander Brace, Carlo Siebenschuh +3

The Large Language Model agent workflow enables the LLM to invoke tool functions to increase the performance on specific scientific domain questions. To tackle large scale of scien…