How to build an AI MVP
The usual MVP advice applies, plus one question it does not cover. An AI MVP has to validate demand and accuracy, and most founders only test the first one.
The standard MVP advice is right and you have read it. Validate the problem before the technology. Pre-sell if you can. Start with the simplest workflow that solves one clear problem. Do the first version by hand behind a form if that is faster than building it.
All of that holds for an AI MVP. There is a second question underneath it that ordinary MVP advice does not cover, because ordinary software does not have this problem.
A normal MVP validates one thing: do people want this. An AI MVP has to validate two, because the second one is whether the thing can be done well enough to be worth wanting.
A landing page with a 6% signup rate tells you there is demand for a tool that reads contracts and flags unusual clauses. It tells you nothing about whether a model can flag them accurately enough that a lawyer would rely on it. Those are separate risks and founders routinely retire the first and discover the second in month four.
Here is the sequence that tests both.
Step 1. Write down what "good enough" means, in a number
Before anything is built. One sentence, with a figure in it.
"It has to catch at least eight of the ten unusual clauses a paralegal would catch, and flag no more than one thing that is not unusual."
That sentence is the most valuable artefact of the whole exercise. It gives you something to test against, it makes the eventual build arguable, and writing it forces the conversation about what your user will actually tolerate, which is a conversation most teams have implicitly and late.
If you cannot write it, that is a finding. It usually means either the task is not well defined or nobody has asked the user what accuracy they need, and both are cheaper to fix now.
Step 2. Do it by hand, and keep the paperwork
Every guide tells you to fake the AI with humans first, and they are right. What they do not tell you is what to keep.
Run twenty or thirty real cases manually. Real inputs from real prospective users, not examples you made up. For each one, save the input and the output a person produced.
That file is your evaluation set. It is the thing you will test every model, every prompt and every future change against, and you have just built it as a side effect of validation. Teams that skip this step build the same file painfully six months later, or more often never build it at all and spend the following year unable to tell whether changes helped.
Doing it by hand also tells you the two things a smoke test cannot:
- How long the task actually takes a person, which is your value estimate, measured rather than assumed.
- Which cases are genuinely hard. Usually a small fraction are, and the shape of that fraction decides the whole build. If the hard ones are rare and low stakes, you have an easy product. If they are rare and high stakes, you have a much harder one.
Step 3. Run the model against the same cases
Now spend an afternoon. Take a current model, write a straightforward prompt, and run your twenty or thirty saved cases through it. Compare against what the person produced.
You are looking for one of three answers:
- It clears your bar already. More common than founders expect. Build the product around it and stop worrying about the model.
- It is close. Then the work is retrieval, prompt structure and handling the specific cases it misses, which is a normal engineering project with a known shape.
- It is not close. This is the answer worth paying for. You have found it in an afternoon rather than in a seed round, and the question becomes whether the product works with a person in the loop, which is frequently a better business anyway.
Do this before the landing page if you can. It costs less than the ad spend.
Step 4. Build the narrowest version that a real user touches
Now the ordinary MVP advice applies again. One workflow, the simplest one, in front of people who have the problem.
Two things specific to AI that are worth doing from the first version:
- Log every input and every output. Not for analytics. For the evaluation set, which should be growing every week from real usage rather than sitting at the thirty cases you started with.
- Make correcting it one click. Every correction is both a better experience and a labelled example. Products that make correction easy accumulate the asset that makes them hard to copy; products that make it hard accumulate nothing.
Step 5. Decide what you are measuring, and measure it before launch
Pick the number the feature is meant to move, and read it before anyone can use the feature. Time to complete the task, proportion of users who finish it, tickets raised about it.
It takes an afternoon and it is the difference between knowing in six weeks whether this worked and having an argument about it.
One limit worth stating
Everything above assumes you can get twenty or thirty real cases before you build. Sometimes you genuinely cannot, because the data sits inside companies that will not share it before they are customers. That is a real constraint and the honest workaround is worse than the real thing: build the narrowest version, get one design partner, and treat their first month as step 2 with the paperwork kept. It is slower and the sample is biased towards one customer, and it is still better than shipping with no idea what good looks like.
The thing that decides this
The founders who get this right are not the ones who picked the better model. They are the ones who wrote down what good enough meant before they started, and kept the file of real cases that let them check.
Everything after that is ordinary product work. If you want the cost side of it, what a build like this takes in engineering weeks is free and splits it by line item. If you are choosing between several AI ideas rather than validating one, start with how to rank them. And before you ship, read what happens when your AI agent gets it wrong, because the failure path is the half that decides whether people keep using it.
If you would rather have somebody run steps 1 to 3 with you in a week, that is the conversation.
If you are deciding what AI belongs in the product before the next raise, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- Is Omarchy the AI-native OS for power users?Omarchy puts coding agents in the operating system rather than in a browser tab. That is the right direction, and it makes one missing piece obvious.
- OpenClaw vs Hermes vs Grok Bot, for companiesThe three agent harnesses compared on the questions a company has to answer rather than a hobbyist. Hosting, credentials, real cost, and who fixes it.
- What is an agent harness?The harness is everything around the model that lets it act on its own. It decides more about whether an agent works in production than the model does.
Get the next one
New writing in Building with AI as it goes up, roughly twice a month. One article per email and nothing else in it.