4 papers
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
Michael J. Zellinger, Matt Thomson
State-of-the-art reasoning LLMs are powerful problem solvers, but they still occasionally make mistakes. However, adopting AI models in risk-sensitive domains often requires error…
Economic Evaluation of LLMs
Michael J. Zellinger, Matt Thomson
Practitioners often navigate LLM performance trade-offs by plotting Pareto frontiers of optimal accuracy-cost trade-offs. However, this approach offers no way to compare between LL…
Rational Tuning of LLM Cascades via Probabilistic Modeling
Michael J. Zellinger, Matt Thomson
Understanding the reliability of large language models (LLMs) has recently garnered significant attention. Given LLMs' propensity to hallucinate, as well as their high sensitivity…
Cost-Saving LLM Cascades with Early Abstention
Michael J. Zellinger, Rex Liu, Matt Thomson
LLM cascades deploy small LLMs to answer most queries, limiting the use of large and expensive LLMs to difficult queries. This approach can significantly reduce costs without impac…