4 papers · 1 filter
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
Daniel R. Jiang, Jalaj Bhandari, Yukai Yang +2
Optimizing large language models (LLMs) for multi-turn conversational outcomes remains a significant challenge, especially in goal-oriented settings like AI marketing or sales agen…
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
Irene Wang, Newsha Ardalani, Mostafa Elhoushi +6
Machine learning solutions are rapidly adopted to enable a variety of key use cases, from conversational AI assistants to scientific discovery. This growing adoption is expected to…
Optimization-Driven Adaptive Experimentation
Ethan Che, Daniel R. Jiang, Hongseok Namkoong +1
Real-world experiments involve batched & delayed feedback, non-stationarity, multiple objectives & constraints, and (often some) personalization. Tailoring adaptive methods to addr…
AExGym: Benchmarks and Environments for Adaptive Experimentation
Jimmy Wang, Ethan Che, Daniel R. Jiang +1
Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for tes…