4 papers
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
Kevin Han, Renfei Zhang, Kathy Wei +3
LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across div…
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
Hamed Mahdavi, Pouria Mahdavinia, Alireza Farhadi +7
State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 o…
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
Hamed Mahdavi, Pouria Mahdavinia, Samira Malek +7
State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 o…
Harnessing Optimization Dynamics for Curvature-Informed Model Merging
Pouria Mahdavinia, Hamed Mahdavi, Niloofar Mireshghallah +1
Model merging is an effective post-training strategy for composing capabilities in large language models without joint retraining. We study this in the supervised fine-tuning (SFT)…