most citedHiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

8 citations · 8 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Self Distillation Fine-Tuning of Protein Language Models Improves Versatility in Protein Design

Amin Tavakoli, Raswanth Murugan, Ozan Gokdemir +3

Supervised fine-tuning (SFT) is a standard approach for adapting large language models to specialized domains, yet its application to protein sequence modeling and protein language…

cs.CL2025

Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models

Ozan Gokdemir, Neil Getty, Robert Underwood +5

As scientific knowledge grows at an unprecedented pace, evaluation benchmarks must evolve to reflect new discoveries and ensure language models are tested on current, diverse liter…

cs.IR20258 cited

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace +21

The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval Augm…

cs.IR2025

AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine

Carlo Siebenschuh, Kyle Hippe, Ozan Gokdemir +10

Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive…

cs.DC2025

Connecting Large Language Model Agent to High Performance Computing Resource

Heng Ma, Alexander Brace, Carlo Siebenschuh +3

The Large Language Model agent workflow enables the LLM to invoke tool functions to increase the performance on specific scientific domain questions. To tackle large scale of scien…