A Feasibility Condition for Confidence-Gated Tiered LLM Inference
Abstract
Running large language models across a mix of tiers promises lower cost and latency by resolving easy queries locally and escalating hard ones upstream. That escalation hinges on a confidence signal the small model emits about its own output. This paper shows the binding constraint inverts: where compute is billed, cost decides; where it is sunk on owned hardware, signal quality decides.
Preprint soon