15 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.CL2026
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
Abel Yagubyan
LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study…
cs.CL2026
How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines
Abel Yagubyan
Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability question remains under-explored: doe…
astro-ph.HE2022★ 15 cited
The Lick Observatory Supernova Search follow-up program: photometry data release of 70 stripped-envelope supernovae
WeiKang Zheng, Benjamin E. Stahl, Thomas de Jaeger +85
We present BVRI and unfiltered Clear light curves of 70 stripped-envelope supernovae (SESNe), observed between 2003 and 2020, from the Lick Observatory Supernova Search (LOSS) foll…