Can there really be too much of a good thing?

Sydney, Australia

Sam Russell, Product Lead

At what point should you stop measuring cloud and AI cost?

Every conversation about technology cost arrives at the same request eventually: can we see this in more detail. It comes from the board, from a product owner defending a margin, from whoever was handed the bill and asked to explain it, and it is always a reasonable thing to ask.

The person asking is usually right in the narrow sense. There is always a more precise answer, although it is almost never sitting there waiting to be read. Getting to it takes tagging that does not exist yet and engineering time nobody has budgeted. What almost never gets asked is whether the answer is worth what it costs to reach, because precision sounds like rigour, and rigour is difficult to argue against.

The maths is never the problem. Every version of a unit cost is the same division, spend over volume, and what changes is which spend and which volume. Divide the whole technology bill by the whole business and you get a number that describes nothing in particular: a shared platform averaged in with a customer-facing product, an AI feature averaged in with a batch job that runs overnight. Divide the spend attributed to one product by that product's own volume, and the same arithmetic produces something of use to finance.

What a unit cost is made of

Take the physical version, because it is the one nobody argues with. A grocer sells 2,000 avocados a month, and getting them onto the shelf costs $4,000, so each one costs $2.00. Inside that $2.00:

  • Fruit: 68c

  • Freight: 30c

  • Labour: 37c picking and packing, 10c shelf stacking, 3c for the stackers' uniforms

  • Premises: 52c of store rent

Six components, not one of them measured on the individual avocado, all of them apportioned across the 2,000. The grocer prices at $3.20 and knows the margin holds at that volume.

They behave differently, which matters more than their sizes. Fruit and freight are variable, two thousand avocados cost twice what a thousand do. Stacking labour and uniforms are fixed until they are not, one more crate does not add to them and a second store does. The rent is joint, shared with the bread and the washing powder, so the 52c is a cost the avocados did not cause. Stop selling them and the fruit and the freight stop. The rent carries on.

The same operation on software

Take a reporting view inside a SaaS application, the kind of screen a customer opens forty times a day. One call touches the application, a gateway, a database query, storage and egress, and a logging pipeline. Say it costs 1.8c. Inside that 1.8c:

  • Compute: 0.4c application, 0.3c gateway

  • Database: 0.5c of query

  • Storage: 0.2c storage, 0.1c egress

  • Observability: 0.3c of logging

Six components again, apportioned the same way.


What happens when you bolt AI on?

Then the product adds an AI assistant to that view, the sort that summarises what the customer is looking at. It does not replace the feature, it sits on top of it, and only some calls invoke it. When one does, the call carries everything above plus 2.1c for the model to produce the summary, 0.4c to find the material it summarises and 0.1c for data brought in from outside. So that call costs 4.4c rather than 1.8c.

If one call in eight invokes the assistant, the average lands at 2.1c, and no call costs 2.1c, they cost 1.8c or they cost 4.4c. The average is correct but it describes nothing that happened.

Which holds identically whether the thing sitting on top is an AI assistant, a batch job or a premium tier. The point is not that AI is expensive, it is that it is different.

Different in how it behaves, most of all. The application, the database and the logging are capacity, bought ahead in blocks and paid for whether the customer opens the screen or not. The assistant is not. It costs nothing until it is asked, and once asked it costs in proportion to what it was asked for, which makes it the fruit. Cloud made a fixed cost partly variable. This makes it properly variable, a cost of sale that moves with volume the way the avocados do, and the first large line on a technology bill that has ever really behaved that way.

Separated out, both figures are enough to negotiate a contract on, enough to price a renewal, and enough to decide where the next increment of investment goes, none of which needs a telemetry programme in front of it. Closing something down asks one thing more, which is knowing which of those costs would actually disappear, because the store rent does not leave with the avocados.

The usual objection is that the volume you divide by is an estimate too, so the whole ratio is soft. It is not. Customers served, transactions processed and calls made are counted by systems that already have to be correct for billing and revenue recognition. The soft side has always been the spend sitting on top, and only in one respect: the total is certain, it arrives on an invoice, and what is estimated is how much of it belongs to this product rather than that one.

What has never had an owner is the question of how precisely that spend needs to be split. Engineering strives for more granularity, aiming for theoretical perfection, because in theory it is optimal. Finance does not know what that costs to produce, and engineering does not consider it. There is no forum in which the cost of measuring is weighed against the value of the answer, so the estate defaults to building either nothing, or everything and never getting there.

The part that costs something

The more precise answer is the correct one. Measuring the cost of each individual request, rather than apportioning a pool across a group of them, is the right method, but it is only marginally better for any figure a decision actually uses. In one measured production workload the median request produced 25 output tokens while the ninety-ninth percentile produced 509, a ratio of roughly twenty to one. Only part of an assisted call moves with that, because the application, the database, the storage and the logging cost much the same whatever the model returns, but on the numbers above it still puts the dearest assisted calls at something like ten times the typical one. So 4.4c describes the typical assisted call accurately and the expensive one not at all. And this is not an argument from inability: OPTIMAZE can resolve cost to the token. The reason to stop short of doing it everywhere is not that it cannot be done.


An average is only a lie when you take it across things that have nothing to do with each other, and once they belong together, further precision costs more than the decision it improves.

Three reasons to stop anyway

The first is arithmetic, a business does not make decisions about single requests. It negotiates a contract covering millions of them, prices a product for a year, decides whether a feature earns its place over a quarter. And across a run of that size the average is not an approximation of anything, because the total it comes from is not an estimate. The invoice is the total. Attribution splits a number that is already known, so however wide the spread between one request and the next,  it cannot move the figure the quarter gets judged on. Precision only pays where the decision can still change the number. Split the shared platform down to the request and the figure gets better while the decision does not, because the capacity was committed before the question was asked. What is still open to a decision is a minority of most bills, so measuring everything puts most of the effort into something with minimal impact. Knowing which part matters, costs a fraction of measuring all of it.

The second is separation, splitting the assisted calls from the plain ones is exactly what attribution derived from how resources, accounts and metadata actually relate produces, with no request instrumented anywhere. The objection is an argument against mixing. It is not an argument against dividing.

The third is detection, because the instrument for finding an outlier is not a measurement programme. Anomalies in OPTIMAZE watches for the request, the workload or the account that does not fit its own pattern, and it does that without pricing every request individually, because detection and measurement are different problems. Detection needs to know what normal looks like and flag the departure. Measurement prices everything whether or not anything is wrong. Building the second to get the first is how a finance team ends up funding a standing engineering programme to answer a question a threshold would have answered.

Robin Cooper's cost management framework holds that the design of a cost system trades the cost of measurement against the cost of errors, and that the optimum sits where the sum of the two is lowest, which is never at maximum accuracy. It also says that optimum shifts with how varied a business is, and on that reasoning a diverse estate belongs closer to full measurement, because the cost of being wrong rises with the number of different things being averaged together. True when measuring more was the only way to be more accurate. It no longer is. What moves is not the stopping point but the resolution of the attribution underneath it.


For most businesses that place is an attributed unit cost, and precision beyond it means paying to know something the contract, the renewal and the investment decision do not turn on. Which makes the practical move the opposite of the instinctive one: the budget that would have funded per-request measurement buys more by making the attribution finer, from a group of products down to the product, the team, the user.

Measuring technology spend accurately enough to act on is the part that pays for itself, and it is enough to quantify what an estate produced rather than treat it as an operating cost to be reduced. Measuring it perfectly is a project, and like any other project the return on it should be quantified before it starts.

That is what OPTIMAZE calls Technology Capital Performance.

If this sounds like something you or your business needs, get in contact with our sales team.

Subscribe to our blog.

Never more than once a week.

Technology Capital Performance is the measurement of whether technology spend is producing a proportionate, attributable return.

© 2026 Optimaze Services Pty Ltd

Technology Capital Performance is the measurement of whether technology spend is producing a proportionate, attributable return.

Technology Capital Performance is the measurement of whether technology spend is producing a proportionate, attributable return.