4 papers
Evaluating Small Language Models for Front-Door Routing: A Harmonized Benchmark and Synthetic-Traffic Experiment
Warren Johnson, Charles Lee
Selecting the appropriate model at inference time -- the routing problem -- requires jointly optimizing output quality, cost, latency, and governance constraints. Existing approach…
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
Warren Johnson
The rapid proliferation of Large Language Models has created an environmental paradox: the very technology that could help solve climate challenges is itself becoming a significant…
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
Warren Johnson
Prompt compression is often evaluated by input-token reduction, but its real deployment impact depends on how compression changes output length and total inference cost. We present…
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
Warren Johnson, Charles Lee
The economics of prompt compression depend not only on reducing input tokens but on how compression changes output length, which is typically priced several times higher. We evalua…