6 papers
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Lei Shi, Anlan Zhang, Rita Lyu +6
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges…
Experimentation Accelerator: Interpretable Insights and Creative Recommendations for A/B Testing with Content-Aware ranking
Zhengmian Hu, Lei Shi, Ritwik Sinha +2
Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and oft…
Dialectics for Artificial Intelligence
Zhengmian Hu
Can artificial intelligence discover, from raw experience and without human supervision, concepts that humans have discovered? One challenge is that human concepts themselves are f…
AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning
Yueqian Lin, Zhengmian Hu, Jayakumar Subramanian +4
Effective human-AI collaboration on complex reasoning tasks requires that users understand and interact with the model's process, not just receive an output. However, the monolithi…
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
Yueqian Lin, Zhengmian Hu, Qinsi Wang +6
We present Voice Evaluation of Reasoning Ability (VERA), a benchmark for evaluating reasoning ability in voice-interactive systems under real-time conversational constraints. VERA…
Cautious Next Token Prediction
Yizhou Wang, Lingzhi Zhang, Yue Bai +7
Next token prediction paradigm has been prevailing for autoregressive models in the era of LLMs. The current default sampling choice for popular LLMs is temperature scaling togethe…