Engineering preview
Prove the cheaper model is good enough. Then change the routing.
Cost analysis tells you which workload is expensive. It does not tell you whether a cheaper model can do that workload. The control plane is the part that answers the second question before anything in production moves.
Where this stands today
- Runs locally: one command starts the web server, one starts the worker.
- 94 tests and an eight-step acceptance demonstration pass on a fresh database.
- No hosted deployment. Storage is a local SQLite file, which is not durable on serverless hosting.
- No live provider evaluation has run yet, so there is no measured quality or cost result to show.
The loop
Four stages, and the fourth is the point.
01
Observe
Every request is recorded with the workflow it belongs to, the model actually used, per-attempt token usage and timing. Prompt and response contents are never stored.
02
Ontology
A workflow's requirements are written down and versioned: how often it must be correct, how often it must decline to answer, and which languages have to clear the bar on their own.
03
Enforce
A model becomes eligible only by passing a reviewed evaluation against that requirement version. A policy then proposes a route, runs in shadow until someone enables it, and can be returned to observation in one action.
04
Observe again
The same ledger records what the change did, including every fallback and retry. Rollback is a mode change, not a redeploy.
How the numbers are kept honest
The spend you are shown is the spend that happened.
A comparison is only worth acting on if the losing side was charged for everything it cost. These are enforced in code and covered by tests, not described in a policy document.
Serving cost includes every billable attempt. A retry that consumed tokens was paid for, and it appears in the total with a breakdown beside it.
An attempt whose token usage the provider did not report is charged at its ceiling, never at zero.
When any attempt cannot be priced, the total is labelled a lower bound and no cost difference or savings figure is reported from it.
A paid evaluation cannot start without an explicit total spending cap, checked before every attempt including retries.
A mock run cannot establish that a model is eligible, however good its numbers look.
An unreviewed answer is never counted as a correct one.
What it will not do on its own
Observation is the default. Enforcement is a decision.
A workflow starts in observation mode and stays there. In shadow mode a policy records the route it would have taken and changes nothing. Routing only changes after someone activates enforcement, and every activation records who did it and why. Returning to observation is one action.
An eligibility decision always carries a reason, and overriding one requires a named person. Nothing is inferred from a model's name or price.
Access
Source is private while the first evaluation is still ahead.
The repository is available on request to design partners. If you run a workload where a cheaper model might be good enough and you need to prove it before switching, that is the conversation worth having.