Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators
Ye Chen, Weining Zhang
Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evalu…
cs.AI2026
UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists
Ye Chen, Weining Zhang
Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapte…
cs.AI2026
Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation
Ye Chen, Weining Zhang
Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weights, or rule-based deferral policies for s…