How to add AI to your SaaS product
A sequence for adding an AI feature to a product that already has customers, and the four things that decide whether it survives its first month.
You have a working product, paying customers, and one or two engineers. Somewhere on the roadmap is a line that says AI, and it has been there for a while because nobody is sure what it means.
Here is the sequence that works for a team that size. It assumes you are adding to something people already use rather than starting from nothing, which changes almost every decision.
Step 1. Pick a task, not a technology
Write the candidate as something a user does today, in their words. "Find the right template for this client" rather than "recommendations". "Explain why this invoice was rejected" rather than "chatbot".
The reason is not style. A task has a frequency you can count and a failure mode you can describe, and a technology has neither, which means a roadmap line saying "AI search" cannot be argued with and a line saying "find the right template" can.
Two questions narrow the list fast:
- Is your user already doing this, badly, inside your product? Look at what people paste into free-text fields, what they export to a spreadsheet, and what your support queue is actually about. Those are tasks your product implied and did not finish.
- Would you notice if the feature were switched off? If nobody would complain, it is a demo rather than a feature. Plenty of AI ships this way and it is the cheapest mistake on this list to avoid.
Step 2. Decide whether it is really a model problem
A meaningful share of what gets scoped as an AI feature is a search problem, a rules problem, or a form with too many fields. Those are cheaper to build, faster, and right every time.
The test: could a person write down the rule in an afternoon? If yes, write the rule. Models are for the cases where the rule is either unknown or has too many exceptions to enumerate, and paying per token to approximate a rule you could have written is a cost that recurs forever.
Where the honest answer is that half the cases are a rule and half need a model, build the rule first and let the model handle the remainder. That is usually a smaller, cheaper and more reliable feature than the version where the model handles everything.
Step 3. Bolt it on rather than rebuilding
Almost nothing about adding AI to an existing product requires touching the existing product. The pattern that works is a separate service that reads through the APIs you already have and writes results back the same way.
Practically that means:
- A new service, deployed separately, so it can fail without taking the product with it.
- Reading your existing data rather than a new pipeline. Whatever the feature needs is almost certainly already in your database, and building a new data flow before you know the feature works is the most common way to spend a quarter on nothing.
- A feature flag in front of it, so it can be turned off for one customer, which you will need in the first fortnight.
You do not need a data science team and you do not need to train anything. Current models handle the overwhelming majority of product features out of the box, and the work is in what surrounds them rather than in the model.
Step 4. Build the evaluation set before the feature
This is the step guides skip, and it is the one that decides whether the feature is still working in six months.
An evaluation set is fifty to two hundred real inputs from your own product with the answer you would accept written next to each one. Not synthetic examples. Real ones, pulled from what your users actually typed.
It takes a day or two to assemble and it buys three things nothing else buys:
- You can tell whether a prompt change made things better or worse, rather than guessing from the three examples you happened to try.
- When the model provider ships a new version, you can find out in an hour whether to move.
- When a customer says it got worse, you have something to check rather than an argument.
Teams without one are not being careless. They are shipping fast, and the cost arrives later as an inability to change anything with confidence.
Step 5. Design the failure path with the happy path
Your feature will be wrong sometimes, and the question is what the user sees when it is.
Decide these before launch, not after the first complaint:
- What it does when it does not know. Saying so is a feature. Guessing confidently is the thing that loses trust and does not come back.
- How the user corrects it. A correction is a fix for that user and a new row for your evaluation set, which is the cheapest data you will ever collect.
- What support sees. They will get the tickets. Give them a way to look at what the feature actually did rather than asking the customer to describe it.
If the feature takes an action rather than producing text, add an approval step for the first version. Propose the refund, let a person confirm it. The approval logs then tell you whether it can be trusted to act on its own, which is a much better basis for that decision than confidence.
Step 6. Check the unit economics before you price it
An AI feature carries a cost per use that the rest of your product does not, and per-seat pricing on a per-use cost breaks quietly, at exactly the point where customers start liking it.
Work out the cost per user per month at your real usage rather than your expected usage, and check what happens at five and ten times that. The tool that does this is free and takes about a minute.
The number that matters is not the token cost. It is the token cost against what that customer pays you, and whether it still leaves you the margin you need.
One limit worth stating
This sequence assumes the feature is a good idea. It has nothing to say about whether it is. A team can execute all six steps well and ship a feature that nobody uses, because the deciding question was answered in step 1 and it was answered by opinion. If you have several candidates and no clear favourite, work out which one is worth building first, because a well-built version of the wrong feature costs more than a badly built version of the right one.
What this costs
Expect the model work itself to be the smaller part. The four things around it, the evaluation set, the failure path, the cost controls, and the documentation for whoever inherits it, are frequently larger than the feature and they are the lines missing from most estimates.
If you want that split into engineering weeks for your specific case, there is a tool for that too, and it gives you a scope you can put in front of a supplier so their quotes are comparable.
Start here this week
Open your support queue and read the last fifty tickets. Group them by what the person was trying to do rather than by what went wrong. The largest group that your product could have answered and did not is your first candidate, and you now have both the frequency and the beginnings of an evaluation set from the same afternoon.
Then read what happens when your AI agent gets it wrong before you write any of it, because that is the half of the build that decides whether the feature is still switched on next quarter. If you would rather have somebody run this with you, that is a 30-minute conversation.
If you are deciding what AI belongs in the product before the next raise, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- Is Omarchy the AI-native OS for power users?Omarchy puts coding agents in the operating system rather than in a browser tab. That is the right direction, and it makes one missing piece obvious.
- OpenClaw vs Hermes vs Grok Bot, for companiesThe three agent harnesses compared on the questions a company has to answer rather than a hobbyist. Hosting, credentials, real cost, and who fixes it.
- What is an agent harness?The harness is everything around the model that lets it act on its own. It decides more about whether an agent works in production than the model does.
Get the next one
New writing in Building with AI as it goes up, roughly twice a month. One article per email and nothing else in it.