Blog / AI unit economics 8 min read

AI unit economics: the hidden cost of coding limits

The AI programs that scale are managed by cost per shipped outcome and available capacity—not by seats purchased or tokens consumed.

Buying an AI coding subscription can feel like buying a fixed block of engineering capacity. The invoice is predictable. The work is not. A limit can arrive in the middle of a change, and the seat price will never show what happened next.

The finding runs against the most comfortable story in AI engineering: buy a subscription and the economics are solved. A subscription makes procurement predictable. It does not tell you what the spend produced, which model carried which part of the work, or what happened when the workflow reached a capacity ceiling.

Provider consoles answer the provider's question: how much did the account consume? Engineering leaders have a different question: what should we keep, fix, or route somewhere else? That decision starts when cost is connected to a shipped outcome.

“A token total cannot tell a CTO whether an AI investment is working. Cost per shipped merge can—and capacity interruptions show where a predictable seat stops being predictable engineering throughput.”

Eric Lam, co-founder and CTO of Codiedev

Four numbers change the decision

  • $8.45Claude Code / merge
    Consumption became an outcome metric. Claude Code recorded $109.79 in API-equivalent usage across 13 merged requests.
  • 15.3hEstimated waking hours to next productive session
    Capacity had an operational shape. Three distinct limit episodes occurred in two days. The figure subtracts an assumed eight-hour sleep window from the elapsed interval.
  • $1.30GLM-5.2 contribution per merge
    Model-level attribution exposed the routing mix. GLM-5.2 contributed $1.30 per merge.
  • $0.19GPT-5.6 Luna contribution per merge
    Fallback economics became visible. GPT-5.6 Luna contributed $0.19 per merge when routed through the alternate model path.

Output tells you what the spend bought

Tokens are a raw material. Seats are a purchasing unit. Neither is the product of an engineering organization. The product is working software that makes it through review and merge. When AI cost is joined to that event, a consumption report becomes a management instrument.

Outcome-linked AI cost30-day telemetry
Abstract three-model cost per merge visualization
Model / ledgerCost per merge
Claude Code · API-equivalent$8.45
GLM-5.2 · metered contribution$1.30
GPT-5.6 Luna · metered contribution$0.19
Cost per merge is shown by model using the captured outcome attribution.

Across the 30-day window, 13 merged requests were associated with captured Claude sessions for the anonymous user. The captured usage carried $109.79 of API-equivalent model value, or $8.45 per merge.

On a flat-price subscription, the equivalent cost for this user was $150 of spend.

For similar work by the anonymous user, GLM-5.2 cost $1.30 per merge and GPT-5.6 Luna cost $0.19 per merge.

For this workflow, Claude Code was the more expensive default on a per-merge basis.

This distinction is where AI unit economics becomes useful. A flat license allocation answers finance's budgeting question. API-equivalent usage answers what captured consumption would cost at list price. Provider-reported spend answers what routed calls cost. Outcome attribution answers what each ledger produced.

The seat price leaves out the capacity tax

A fixed-price seat is attractive because the maximum cash outlay is known. The hidden variable is available capacity. When usage limits interrupt the work, the economic question changes from “what did the seat cost?” to “what work could still use the seat when it was needed?”

The anonymous user reached three distinct capacity limits between the afternoon of August 4 and the afternoon of August 5. New sessions were attempted after the first and third episodes, but no captured productive tool work followed in those sessions. The next captured productive AI session began at 3:06 p.m. on August 6, 23.3 elapsed hours after the third limit—or 15.3 estimated waking hours after subtracting an assumed eight-hour sleep window.

Most license dashboards stop at assignment and activity. They can tell you that a seat exists and that the account used it. They cannot connect the limit event to the next productive session, show whether another model carried the work, or calculate the cost of the work that ultimately merged. That missing join is the capacity tax hiding inside a fixed-price plan.

The answer is routing, not replacement

Engineering leaders do not need another ideological argument about closed versus open models. They need a routing policy. Keep premium subscriptions where they produce high-value work. Add metered capacity where a task can be completed at the required quality and where a limit would otherwise stop the AI-assisted workflow. Then measure both paths against the same shipped outcome.

Provider fallback after a limitroute, then compare outcomes
Abstract fallback routing visualization showing a primary coding lane branching to alternate model providers
GLM-5.2lower-cost alternate route
GPT-5.6 Lunaalternate route to compare on outcome
A capacity limit does not have to end the workflow. The engineer can route the next task through an available provider, then use model-level performance and cost data to decide which alternate path is worth keeping.

When Claude Code reaches a capacity limit, the engineer can route the next task through an available provider with a broader model catalog. GLM-5.2 and GPT-5.6 Luna become alternate paths—not automatic replacements. The performance view then shows which route produces dependable merges at the lowest cost for each kind of work.

The routing decision cannot be made from price alone. Quality, review burden, retry cost, context handling, and merge success all affect the real unit cost. A cheap run that fails is waste. A premium run that finishes the hardest task in one pass can be the lowest-cost option. The governing metric is cost per successful shipped outcome, segmented by model and work type.

A CTO's AI unit-economics playbook

The operating model starts with one rule: every AI cost needs an outcome. That outcome can be a merged request, a successful agent run, a resolved incident, or a customer-facing product event. Without one, the organization can report consumption but cannot govern the investment.

Separate the ledgers. Track actual subscription cash, API-equivalent usage, and provider-reported metered spend as different fields. A subscription should be allocated to the developer and tool that owns it, not spread across the entire company. Metered calls should retain their exact provider and model.

Attach work. Link sessions to branches, commits, and merged requests. Preserve model attribution instead of forcing every outcome to have one model owner. The three metered-model merges in this analysis also had Claude activity; the useful truth is the composition, not a winner-take-all label.

Measure capacity. Collapse repeated limit messages into distinct episodes, record the next productive AI activity, and identify whether another model picked up the work. Limits are not merely user-experience events. They are a constraint on the engineering capacity the company purchased.

Make the decision visible. A useful executive view answers four questions in one screen: What did we pay? What shipped? Where did capacity fail? Which subscriptions and models should we keep, fix, or reroute? That is the difference between an AI adoption dashboard and AI engineering mission control.

Reading the signal on your own team

The mechanism is general: cost becomes governable when it is attached to a real outcome. Start with your highest-usage workflows, because that is where both the value and the capacity constraints show first. For each tool and workflow, calculate cash allocated per linked merge, API-equivalent usage per linked merge, provider-reported spend per linked merge, limit episodes, and the interval to the next productive AI event.

Then compare work, not people. The purpose is not to rank engineers by consumption. It is to discover which combinations of model, task, and purchasing plan produce dependable shipped work.

Codiedev connects AI activity, model cost, capacity events, and shipped engineering outcomes so leaders can manage the portfolio from real work. The result is a decision layer above provider consoles: not who used the most AI, but which investments are creating output and where the next dollar of capacity should go.

About Codiedev

Codiedev is mission control for AI engineering. It connects AI adoption, model spend, capacity, and shipped outcomes from real engineering work, giving leaders the control layer to drive performance and ROI across every model their teams use.

Mission control for AI engineering

Grab the controls.

See your AI cost per shipped outcome, where capacity is being lost, and which models your team should keep, fix, or reroute.

Book a 20 min demo