Skip to main content
Valfiguer
Laptop showing an analytics dashboard, charts and metrics reflected on a glossy table
Photo · Unsplash
Responsible AI

What it really costs to put AI inside your product

Nicolás Valfiguer
Nicolás ValfiguerFounder, Valfiguer
Sep 8, 20268 min read

There is an awkward distance between showing an AI demo and invoicing a product with AI inside it. The demo is paid for by your curiosity. The product is paid for by a customer every month, and between the two appears a line in the accounts that almost nobody looks at until it hurts: what it costs you to serve that customer.

Stack Overflow published the second part of its analysis on the economics of agents at scale this week, and GitHub described how it made its own assisted coding cheaper without lowering quality. Both pieces circle the same thing from different sides: cost per token stopped being an infrastructure detail and became a product decision.

We run Projekt Republic, an operations and finance SaaS we also use to run ourselves. That puts us on both sides of the problem: we write the invoice and we receive one. What follows is how the cost looks from there.

Not infrastructure — cost of goods sold

The most expensive framing error is filing AI in the same mental box as the server. A server is a fixed cost: you pay the same with ten customers or two hundred, and every new customer improves your margin. AI does not work that way. Every user who opens the assistant spends tokens you would not have spent had they stayed out.

Put in the language of the accounts: AI is a variable cost, and variable costs belong in cost of goods sold, not overheads. The moment you move that box, your product's gross margin changes, and so does everything hanging off it — price, break-even, how much you can spend to win a customer.

Three questions before writing a line

Before putting a model behind a feature you intend to charge for, three numbers are worth having. Not tidy estimates: numbers.

  • What does a typical interaction consume? Not the best or the worst — the median, measured on real use rather than your desk test, which is always shorter and cleaner.
  • How many interactions a month does a user who gets value out of it make? The heavy user defines your cost, not the average. The average is dragged down by people who opened the feature once.
  • What happens when the model gets it wrong? A retry is another call. A three-step flow whose middle step fails half the time costs considerably more than the sum of its steps.

The third is the one most often skipped and the one that surprises hardest. An agent's cost is not the cost of its happy path.

The biggest model is rarely the right one

The instinct to wire the most capable model into everything is understandable and expensive. Most tasks inside a product are not open-ended: classify an incoming email, pull four fields off an invoice, summarise a thread. They have a shape, and a small model fitted to that shape does them just as well for a fraction.

GitHub's piece lands in the same place: the saving did not come from cutting quality, it came from no longer using the big hammer on small nails. The practical consequence is that routing between models is an architecture decision, not a later optimisation. Leave it until the bill hurts and the expensive model is already wired into twenty places.

What we measure

At Projekt Republic the operational question is not "how much did we spend on tokens this month". It is how much per organisation, because that is the unit that invoices. An aggregate tells you something went up; a per-customer figure tells you which one.

A second thing falls out of that, and we did not expect it: the distribution is very uneven. A handful of intensive users generate most of the consumption, and they are almost always the ones getting the most value. That is not a problem to trim, it is pricing information. A flat plan over a cost that lopsided quietly moves margin from light users to heavy ones, and sooner or later somebody notices.

A variable cost you cannot split by customer is not a cost. It is a deferred surprise.

What does not change

None of this is new. It is the discipline any business with a marginal cost has always needed: know what one more customer costs you and charge accordingly. What changed is that it now applies to software too, after two decades of the comfort of a marginal cost near zero.

The industry is meeting this at once. While compute spending keeps climbing — the CEO of Nvidia partner Iren says demand may never be sated — whoever builds product on top has to decide how much of it to absorb and how much to pass on.

Our position is simple: AI goes into a product when we can say what serving it costs, not when it looks good in the demo. That is not caution. It is the only way the feature still exists in two years.

Valfiguer

We build and scale technology companies.

Valfiguer Notes

A weekly letter on building and scaling technology companies: what works, what doesn't, and what it cost us to find out.

Subscriptions aren't open yet, so we're storing no addresses at all — we'd rather tell you than pretend you're on the list. Write to us and we'll let you know the day it opens.

Tell me when it opens

© 2026 Valfiguer. All rights reserved.

VALFIGUER LLC — 407 Lincoln Rd STE 708, Miami Beach, FL 33139