3 papers
cs.LG2026
Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study
Kaihua Ding
Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We…
cs.AI2026
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Kaihua Ding
LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-exper…
stat.ML2025
Variance-Bounded Evaluation of Entity-Centric AI Systems Without Ground Truth: Theory and Measurement
Kaihua Ding
Reliable evaluation of AI systems remains a fundamental challenge when ground truth labels are unavailable, particularly for systems generating natural language outputs like AI cha…