activity
20242026
collaborators

7 papers

q-bio.QM2026

Scalable Agentic Reasoning for Designing Biologics Targeting Intrinsically Disordered Proteins

Matthew Sinclair, Moeen Meigooni, Archit Vasan +14

Intrinsically disordered proteins (IDPs) represent crucial therapeutic targets due to their significant role in disease -- approximately 80\% of cancer-related proteins contain lon…

cs.LG2026

LSHBloom: Memory-efficient, Extreme-scale Document Deduplication

Arham Khan, Robert Underwood, Carlo Siebenschuh +7

Contemporary large language model (LLM) training pipelines require the assembly of internet-scale databases full of text data from a variety of sources (e.g., web, academic, and pu…

cs.DC2025

Experiences with Model Context Protocol Servers for Science and High Performance Computing

Haochen Pan, Ryan Chard, Reid Mello +12

Large language model (LLM)-powered agents are increasingly used to plan and execute scientific workflows, yet most research cyberinfrastructure (CI) exposes heterogeneous APIs and…

cs.IR2025

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace +21

The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval Augm…

cs.IR2025

AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine

Carlo Siebenschuh, Kyle Hippe, Ozan Gokdemir +10

Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive…

cs.DC2025

Connecting Large Language Model Agent to High Performance Computing Resource

Heng Ma, Alexander Brace, Carlo Siebenschuh +3

The Large Language Model agent workflow enables the LLM to invoke tool functions to increase the performance on specific scientific domain questions. To tackle large scale of scien…