1 citations · 1 across the 9 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
CT: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees
S M Rafiuddin, Atriya Sen
Sentiment in social-media threads does not only vary across posts; it shifts as users react to claims, corrections, evidence, and hostility within a branching reply tree. We study…
cs.LG2026
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Stephane Collot, Colin Fraser, Justin Zhao +3
Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations…