3 papers
cs.LG2026
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Lei Shi, Anlan Zhang, Rita Lyu +6
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges…
cs.AI2026
Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking
Zhengmian Hu, Lei Shi, Ritwik Sinha +2
Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and oft…
cs.AI2026
QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation
Rita Qiuran Lyu, Michelle Manqiao Wang, Lei Shi
User queries in real-world retrieval are often non-faithful (noisy, incomplete, or distorted), causing retrievers to fail when key semantics are missing. We formalize this as retri…