Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
Warren Johnson
The rapid proliferation of Large Language Models has created an environmental paradox: the very technology that could help solve climate challenges is itself becoming a significant…
cs.CL2026
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
Warren Johnson
Prompt compression is often evaluated by input-token reduction, but its real deployment impact depends on how compression changes output length and total inference cost. We present…
cs.CL2026
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
Warren Johnson, Charles Lee
The economics of prompt compression depend not only on reducing input tokens but on how compression changes output length, which is typically priced several times higher. We evalua…