2 citations · 2 across the 3 of their papers we have counts for
5 papers
Latency and Token-Aware Test-Time Compute
Jenny Y. Huang, Mehul Damani, Yousef El-Kurdi +2
Inference-time scaling has emerged as a powerful way to improve large language model (LLM) performance by generating multiple candidate responses and selecting among them. However,…
Optimal Policy Minimum Bayesian Risk
Ramón Fernandez Astudillo, Md Arafat Sultan, Aashka Trivedi +4
Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniq…
Latent Principle Discovery for Language Model Self-Improvement
Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo
When language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. Howev…
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
Young-Suk Lee, Chulaka Gunasekara, Danish Contractor +2
We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxo…
The Future of Open Human Feedback
Shachar Don-Yehiya, Ben Burtenshaw, Ramon Fernandez Astudillo +17
Human feedback on conversations with language language models (LLMs) is central to how these systems learn about the world, improve their capabilities, and are steered toward desir…