Why Enterprises Pay the Frontier Premium REV 2.1.5

September 2026 update · five mechanisms behind the closed-vs-open procurement decision · modelled economics

⤓ Download PDF
Token price is the wrong denominator. Cost per unit of output is (1 + f) / m — model spend is additive, throughput is a divisor. At a revised 2.4–5.8% of loaded labour, the frontier needs only a 4.0% throughput edge over open weights to break even — a bar that pN compounding clears at any horizon N ≥ 2. The genuine overpayment isn't choosing proprietary; it's routing everything to a flagship max-effort tier.
The single largest driver, and the one invisible on single-turn benchmarks. Success on an N-step agentic task approximates pN. A 2-point per-step reliability edge is nearly invisible in a benchmark table and decisive by step 50.
Task completion probability vs agentic horizon
Y = P(all steps succeed), % · X = steps in the task (log scale)
Per-step reliability10 steps25 steps 50 steps100 steps
Why this dominates
Moving per-step reliability from 97% to 99% — a two-point delta a composite index would round away — is a 2.8× swing in completion at N=50 and 7.7× at N=100. That is the boundary between "the agent ships the PR" and "an engineer babysits and re-drives." Read against driver 2: the completion-ratio gain from that edge is 2.1% at N=1, 4.2% at N=2, 22.6% at N=10 — it clears the 4.0% break-even bar from N=2 onward, by a widening margin. The bridge from completion probability to m is accepted output (limit 4 in driver 2): a task that fails to complete ships no feature and adds nothing to m.
REVISEDREFRAMED
Cost per unit of output is (1 + f) / m, where f is model spend as a fraction of loaded labour and m is the throughput multiplier. f is additive; m is a divisor. A small numerator penalty buys a large denominator gain, which is why the frontier wins while spending several times the tokens.
Reading the formula
Worked example · one engineer, one year
Illustrative units of completed work — not a measurement
What the example shows
Why the asymmetry is the whole argument
Elasticity of unit cost with respect to each input
InputElasticityAt f = 5%Reading
Leverage
Scenario comparison
Scenariof · spendm · throughput Unit costvs baseline
Throughput edge the frontier must deliver to break even
Y = required advantage over open weights, % · X = frontier model spend as % of loaded labour · each curve is a different open-weights spend baseline · shaded band = revised realistic spend
Break-even in closed form
What the identity cannot see
Revised spend band — the earlier 1–2% was anchored too low
A saturated agentic engineer running parallel sessions lands nearer $500–$1,200/seat/month, or 2.4–5.8% of $250k loaded labour, with a fat right tail (MED — inferred from agentic usage patterns, not a surveyed figure). Raising f does not weaken the case: at 5% frontier spend against 1% open weights, break-even is a 4.0% throughput edge. Even at 15% spend it is only 13.9%. The token-price argument was never load-bearing.
Jevons drift — and where the frame breaks
As inference gets cheaper per unit of work, consumption rises faster than unit cost falls, so total spend goes up. The 2.4–5.8% figure will drift upward, and that drift is a health signal rather than a cost problem. Past some point "tooling as a percentage of labour" stops being the right accounting frame entirely: compute becomes a factor of production alongside labour rather than an overhead line charged against it. That is when the procurement conversation changes shape.
NEW IN THIS EDITION
A throughput multiplier can be banked two ways: hold output and cut headcount, or hold headcount and expand output. They are identical on unit cost and completely different on total value. Illustration below: a 100-engineer team at m = 1.45, f = 5%.
Two ways to bank the same multiplier
Indexed to baseline = 100 · cost, output, and cost per unit of output
Mode A
Reduce headcount
Value is capped at the labour line — you recover at most the salaries eliminated, and the gain terminates. Note that f itself does not move: N/m engineers at the same per-seat spend scale numerator and denominator together. What shrinks is the absolute AI budget, which can make a high-leverage line item look small and invite under-investment — a perception risk, not arithmetic.
Forced when: demand is the binding constraint
Mode B
Hold headcount, expand output
Value is bounded by opportunity set, not by cost. Unbounded upside where a project backlog exists — additional revenue-generating work rather than a one-time cost recovery.
Available when: engineering capacity is the constraint
The choice is a diagnostic, not a preference
Both modes land at 72.4 on cost per unit of output — the arithmetic cannot distinguish them. What separates them is which constraint actually binds. Most software organisations claim engineering capacity is their limit, often accurately; firms that reflexively take Mode A are usually revealing that their real constraint was never headcount. That is a strategy disclosure, not a procurement decision.
"Open weights reach ~87% of proprietary" implies linear value transfer. On threshold-gated work — the task either clears the bar or it doesn't — delivered value is sigmoid, and the last few index points sit on the knee.
Delivered value vs model capability
Illustrative: assumed linear transfer vs threshold-gated response · knee at ~88
The composite-average trap, priced
On this curve 87% of frontier capability delivers ~38.9% of the value, not 87%. The benchmark isn't wrong — a mean over ten evals is simply the wrong functional form for work gated on a pass threshold. The easy 80% of tasks were automatable two model generations ago; the residue is exactly what sits on the knee. Illustrative model — LOW confidence on curvature, HIGH on direction.
Open weights are free; serving them is not. Fixed cost is amortised over tokens actually produced, so effective unit cost is governed by utilisation — and enterprise diurnal load runs 5–10× peak-to-trough.
Effective cost per 1M output tokens vs fleet utilisation
Modelled: 4× 8×H200 nodes + 3 ops FTE · shaded band = realistic enterprise utilisation
Model assumptions — load-bearing, tagged
Infra $911k/yr (4 nodes × $26/hr, 8×H200 rented — MED) plus ops $900k/yr (3 FTE × $300k loaded — MED) = $1.81M/yr fixed. Peak capacity 5,000 output tok/s per node for a ~100B-class MoE under heavy batching — LOW confidence; this is the swing variable. Halve throughput and every crossover doubles. The qualitative result is robust to that: self-hosting beats a flagship price almost immediately and a cheap open-weights API only at high sustained load.
Eight drivers, ordered by how much of the premium they explain, with basis and confidence on each.
DriverMechanismBasisConf
Lane 1 · Open weights
Bounded, high-volume, private
Classification, extraction, routing, embedding. Error-tolerant, short-horizon, and sustained enough to clear the utilisation bar. The open-weights case holds cleanly here.
Also: narrow fine-tunes that beat general frontier on one task
Lane 2 · Mid-tier proprietary
General reasoning, good-enough
Where the dominated-flagship trap lives. Cost-tuned proprietary now out-values open weights on both axes — same index, a fifth the cost.
Trap: max-effort by default on work a mid tier clears
Lane 3 · Frontier
Long-horizon, calibration-critical
Agentic chains, must-not-hallucinate, threshold-gated. pN and the sigmoid knee are both load-bearing. Token cost is the small line item.
Plus: indemnity, compliance, harness — often dispositive
Net
Enterprises don't pay the premium because they're bad at arithmetic — they pay it because the arithmetic they're doing has a different denominator, and because f enters additively while m divides. Where open weights genuinely win is a real and sizeable quadrant, but it's bounded work at sustained load, not general reasoning. The overpayment that actually exists is failure to segment traffic, and it shows up inside the proprietary cohort as often as across the closed/open line.
Everything in this document that is a choice rather than a measurement, stated plainly, plus what changed in this edition and why.
Change log
Notes & clarifications