Most production AI cost never shows up on a model bill: orchestration, retrieval, retries, observability. And agents multiply it — a single autonomous task can burn 50–500× the tokens of a chat. Here is where the money actually goes, and how governed execution changes the math.
Sources: third-party industry research (Zylos Research 2026 on token budgets; FinOps analyses from LLM CFO, Addepto, FutureAGI) — not Kernos customer data. The patterns they describe are structural.
The agent loop itself — planning, tool calls, result parsing, retries — is the 72%. Every retry re-reads the same context at full price.
Routing a summarization task to the same frontier model that handles your hardest reasoning is the default outcome of a single-provider setup — and it triples the bill for no quality gain.
Agents re-send the full document, schema, and history on every step. Without deliberate context discipline, you pay for the same tokens hundreds of times per task.
A prompt or model change silently breaks a workflow. Nobody notices until the output is wrong in production — then you pay to re-run everything it touched.
The most expensive token is the one after a bad write: reconciliation, rollbacks, human cleanup across systems. An uncaught wrong action costs more than a month of inference.
If token spend doesn't map to teams, features, or actions, you can't cut what you can't see. Most stacks export one aggregated invoice and stop there.
The same mechanism that makes agents safe also makes them cheap to run — because the expensive failure modes are all failures of control.
Every write-side action is a staged proposal evaluated against your rules before it executes. The wrong write never reaches SAP — so you never pay the recovery tax that dwarfs the token bill.
Each agent action carries its full context onto the audit chain — which objects, which rules, which approval. Token budgets are priced per action tier, so spend corresponds to business activity you can inspect, not an opaque meter.
Built-in evals run your workflows against known scenarios before changes ship. A broken prompt fails a test, not a production payment run.
Bring your own endpoints, choose per-task models — local or private. Route summarization to a small model and spend frontier tokens only where reasoning demands it. No provider lock, no single-invoice cliff.
Minority. Industry analyses consistently put the majority of production AI cost outside the model invoice — orchestration, retrieval, retries, and observability infrastructure. One 2026 analysis (Zylos Research) puts it at 72% outside the invoice. Your model bill is where the spend is visible, not where it ends.
An agent task is a loop: plan, call tools, read results, retry, reflect. Research on token budgeting records autonomous tasks reaching 5 million tokens — 50 to 500× a chat interaction's footprint. The loop is the product; the cost is structural, which is why control mechanisms are the lever that matters.
Three mechanisms: staged execution catches bad actions before they run (avoiding recovery costs that dwarf inference); evals catch regressions before production; and model access is pluggable, so each task routes to a price-appropriate model. We don't claim to lower your model invoice — we claim to remove the cost categories around it.
By action tokens, not seats: Starter $500/month (3 agents, 10K tokens), Professional $2,000/month (10 agents, 100K tokens), Enterprise custom. Every action is staged and audited, so token spend maps to business actions you can inspect.
No. FinOps discipline — attribution, budgeting, routing — works on any stack. What Kernos adds is the control plane: the same staged-execution and audit chain that makes agents safe also makes their spend attributable and their failures cheap. If you already have FinOps tooling and no agent writes to systems of record, you may not need us yet.
Bring one agent workflow. We'll map where its tokens actually go — plan, retrieval, retries, recovery — and what a governed version would cost. No slideware.