collaborators

5 papers

cs.CL2026

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

Stefan Larson, Attila Nagy, Sam Desai +8

RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overla…

cs.LG2026

Quantifying Explanation Quality in Graph Neural Networks using Out-of-Distribution Generalization

Ding Zhang, Siddharth Betala, Chirag Agarwal

Evaluating the quality of post-hoc explanations for Graph Neural Networks (GNNs) remains a significant challenge. While recent years have seen an increasing development of explaina…

cs.LG2026

LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models

Siddharth Betala, Samuel P. Gleason, Ali Ramlaoui +12

Generative machine learning (ML) models hold great promise for accelerating materials discovery through the inverse design of inorganic crystals, enabling an unprecedented explorat…

cs.CL2025

A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation

Siddharth Betala, Kushan Raj, Vipul Betala +1

In this paper, we describe our system under the team name BLEU Monday for the English-to-Indic Multimodal Translation Task at WAT 2025. We participate in the text-only translation…

cs.DL2025

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

Magdalena Lederbauer, Siddharth Betala, Xiyao Li +16

The development of synthesis procedures remains a fundamental challenge in materials discovery, with procedural knowledge scattered across decades of scientific literature in unstr…