2 papers
cs.LG2025
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
Benjamin Feuer, Chiung-Yi Tseng, Astitwa Sarthak Lathe +2
LLM-judged benchmarks are increasingly used to evaluate complex model behaviors, yet their design introduces failure modes absent in conventional ground-truth based benchmarks. We…
cs.LG2019
Self-attention based BiLSTM-CNN classifier for the prediction of ischemic and non-ischemic cardiomyopathy
Kavita Dubey, Anant Agarwal, Astitwa Sarthak Lathe +2
Heart Failure is a major component of healthcare expenditure and a leading cause of mortality worldwide. Despite higher inter-rater variability, endomyocardial biopsy (EMB) is stil…