Startup Launch & Venture Creation Startup Launch & Venture Creation

Build an AI Product That Survives Contact With Production

The demo works because you chose the inputs. Production supplies its own — along with concurrency, latency budgets, a cost line that eats your margin, and a model that quietly changes underneath you. Here is what closes that gap. An AI product survives production when four things exist that a demo never needs: an evaluation […]

Build an AI Product That Survives Contact With Production

The demo works because you chose the inputs. Production supplies its own — along with concurrency, latency budgets, a cost line that eats your margin, and a model that quietly changes underneath you. Here is what closes that gap.

An AI product survives production when four things exist that a demo never needs: an evaluation set that tells you whether a change made things better, a defined behaviour for being wrong, a cost per action that fits inside your gross margin, and a latency budget you have actually measured. None of these are model problems. All of them are engineering and measurement problems, and they are the reason most AI projects stall.

The failure rates bear this out. RAND interviewed 65 experienced data scientists and engineers — 50 from more than 50 organisations and 15 academics — and found that more than 80% of AI projects fail, roughly twice the rate of IT projects that do not involve AI.[1] Their five root causes are almost entirely organisational and infrastructural: misunderstood problems, insufficient data, technology chosen ahead of the problem, underinvestment in infrastructure, and problems beyond what the technology can currently do.

Why demos mislead so reliably

A demo is a curated sample of one. You picked the inputs, you were watching, and when something went wrong you ran it again. Every one of those conditions disappears at launch.

The last row is the one teams underestimate. Conventional software is deterministic: it does the same thing on Tuesday that it did on Monday. An AI product sits on a model that gets deprecated, a prompt someone edits, and data that drifts — and none of those announce themselves. Without measurement, quality changes are invisible until a customer reports them.

In conventional software you ship and it stays shipped. In an AI product, standing still requires active work.

Build the evaluation set before you build the feature

This is the single highest-leverage thing you can do, and it is routinely deferred because it feels like overhead. It is not overhead. It is the instrument that makes every subsequent decision cheap.

A golden set of 100–300 real inputs with acceptable answers costs a few days to assemble and permanently changes how the team argues. Without it, “did that prompt change help?” is settled by whoever is most senior or most recently impressed. With it, the question takes four minutes.

Three details matter more than the tooling you choose:

  • Collect inputs, do not invent them. Invented test cases encode your assumptions about how people will use the product. Real ones contain the typos, the pasted screenshots and the questions you did not anticipate.
  • Read the failures individually. An average score tells you the temperature; the failures tell you the disease. Every useful improvement we have seen came from someone reading fifty bad outputs in a row.
  • Fix the class, not the example. If one input failed, a hundred like it will. Patching the single case makes the score go up and the product stay broken.

Decide what happens when it is wrong

It will be wrong. The question is not how to prevent that but what the product does about it — and that is a product decision, not a technical one.

BehaviourUse whenThe cost
Refuse and say soConfidence is low and a wrong answer is expensiveFeels less capable; frustrating if overused
Escalate to a humanThe stakes are high and volume is manageableReal staffing cost; sets a response-time expectation
Fall back to a simpler methodA deterministic route exists, even a worse oneTwo systems to maintain
Show the workingThe user can verify faster than they could do itOnly works if sources are genuinely checkable
Ask a clarifying questionThe input is genuinely ambiguousAdds a turn; irritating when overused

The failure mode to avoid is confident fabrication with no signal to the user. It is worse than refusing, because it transfers the checking burden to someone who does not know they need to check — and once a user has been burned that way, they stop trusting the correct answers too.

Work out the cost per action before you set a price

Inference is a variable cost that scales with usage, which makes an AI product structurally more like a marketplace than like conventional SaaS. Founders who price on SaaS instincts discover this in month two.

Note what the retry multiplier does. A feature that reruns on failure, falls back to a larger model, or gets a second pass for quality does not cost what a single call costs. In our experience a realistic loaded figure is 1.2 to 1.6 times the naive per-call calculation, before caching.

The levers, in the order they usually pay off: shorten the context you send, route easy cases to a cheaper model, raise your cache hit rate, and reduce retries by fixing the underlying failure class. Switching provider is usually the least effective of these and the first one teams try.

Latency is a product feature

Users tolerate slow if they know why and can see progress. They do not tolerate an unexplained pause. Three practical rules:

  • Budget per step, not per request. A chain of four calls at two seconds each is an eight-second wait, and nobody planned it.
  • Stream anything conversational. First token time matters more to perceived speed than total time.
  • Design for the slowest 5%. Median latency is a comfortable number that describes an experience nobody has. Watch the tail, because that is where abandonment happens.

The statistic everyone is quoting this year

You will have seen the claim that 95% of generative-AI pilots fail. It comes from The GenAI Divide: State of AI in Business 2025, produced by MIT’s NANDA initiative in August 2025, based on 150 interviews with leaders, a survey of 350 employees and analysis of 300 public deployments. The precise finding is narrower than the headline: about 5% of pilots achieved rapid revenue acceleration, while most delivered little measurable impact on profit and loss.[2]

Two things are worth taking from it rather than the number itself. The report attributes failure to a “learning gap” — generic tools that do not adapt to a specific organisation’s workflow — rather than to model quality. And it found that buying a specialised vendor solution succeeded roughly 67% of the time against a far lower rate for internal builds.

For a founder, that is an encouraging finding rather than a discouraging one: it says the winning position is a narrow product that fits one workflow properly, which is exactly what a startup can build and an enterprise IT department usually cannot.

What to build first

An ordered list, and the order matters more than the contents.

  1. The evaluation set. Before the feature. Yes, really.
  2. One narrow workflow, end to end. Not a general assistant. A specific job, done completely, for one kind of user.
  3. Logging of every input and output. With consent and retention rules agreed up front — you cannot improve what you did not keep.
  4. The wrong-answer behaviour. Chosen deliberately from the table above, not left as whatever the model happens to do.
  5. Cost and latency instrumentation. Per action, visible on a dashboard, from day one.
  6. A model-swap seam. One layer where the provider is configured, so switching is an afternoon rather than a quarter.

Notice that only item two is what a user would call a feature. That ratio is roughly right, and it is why AI builds carry a higher infrastructure share of budget than conventional software — commonly 15–20% rather than the usual 5–8%.

Six mistakes we see repeatedly

  • Choosing the technology before the problem. RAND names this explicitly as a root cause. “We should use agents” is not a product decision.
  • No evaluation until after launch. The first month is then spent arguing about whether anything improved.
  • Optimising the average. Users experience individual outputs, not means. A good average with a bad tail is a bad product.
  • Ignoring the data question until it is urgent. Rights, consent, retention and residency shape architecture. Retrofitting them is expensive.
  • Building a general assistant. Broad scope makes evaluation impossible and value hard to prove. Narrow beats capable at this stage, every time.
  • Treating the model as the product. The model is a component available to your competitors on the same terms. The workflow, the data and the trust design are the product.