2 papers
cs.AI2026
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
Emma Yanyang Kong, JJ Tan, Ishan Gupta +8
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accel…
stat.CO2026
A Human-Augmenting Agentic Workflow for Observational Causal Inference
Winston Chou, Adrien Alexandre, Lars Olds +2
Data analysis agents are becoming increasingly common tools for applied and scientific research. Yet, for highly specialized tasks such as Observational Causal Inference (OCI), hum…