1 citations · 1 across the 2 of their papers we have counts for
4 papers
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
Sasi Kiran Gaddipati, Diyana Muhammed, Farhana Keya +2
Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks have failed to keep pace. To…
From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry
Aritra Roy, Kevin Shen, Andrew MacBride +350
Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broa…
SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
Diyana Muhammed, Giusy Giulia Tuccari, Gollam Rabby +2
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse applications, from open-domain question answering to scientific writing, medical decision supp…
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
Gollam Rabby, Diyana Muhammed, Prasenjit Mitra +1
Scientific hypothesis generation is a fundamentally challenging task in research, requiring the synthesis of novel and empirically grounded insights. Traditional approaches rely on…