Free, nothing hidden
The AI due diligence questions that are not on a generic checklist.
16 questions to put to a company claiming AI, written to be read aloud on a call. Each one comes with what a strong answer sounds like and what should worry you, which is the part every template leaves out.
The short version
The short version
A generic due diligence checklist covers finance, legal, HR and market sizing, and tells you nothing about whether the company has built anything. The technical questions that matter are narrower: what a competitor would get wrong if they rebuilt this next month, how the team knows a change made the product better rather than worse, what one customer costs in inference against what they pay, and what happens when the system is wrong. Ask those and the answers separate a real system from a thin wrapper within about twenty minutes. What makes the difference is not having the questions, it is being able to tell a strong answer from a confident one.
Nothing marked yet. Work down the list on the call and mark as you go.
Is any of this theirs?
The first thing to establish, because if the answer is no then most of the rest of the diligence is about a company that can be copied in a fortnight.
“If someone rebuilt this on the same model API next month, what would they get wrong?”
What it tests. Whether there is anything here beyond the prompt. It is the single most useful question on this list and it takes ten seconds to ask.
Sounds like
A specific, concrete answer: the retrieval and how it was tuned, the evaluation set built from years of their own data, the workflow the product sits inside, the integrations that took months of somebody else’s cooperation. A founder who has thought about this answers immediately.
Worry if
A pause. Then either "our prompts are very sophisticated", or a jump to something that is not technical at all, like the brand or the team. Prompts are not defensible. They can be extracted from the product by a competent person in an afternoon.
“What data do you have that a competitor could not buy or scrape?”
What it tests. Whether the data moat is real. Almost every AI pitch claims one and most of the claims describe data that is public, purchasable, or their customers’ property rather than theirs.
Sounds like
Data generated by their own product being used: corrections users made, outcomes that followed a recommendation, labels that came from real work rather than from an annotation contract. It compounds, and they can show it growing.
Worry if
"We have access to a lot of industry data." Access is not ownership. Also treat with care any answer where the data belongs to customers and the contract does not clearly grant the right to learn from it, because that is a moat that can be revoked in a renewal.
“What happens to the company if your model provider triples the price or shuts the model down?”
What it tests. Concentration risk on a supplier they do not control, and whether they have ever thought about it.
Sounds like
They have run their evaluation set against at least two providers, know roughly what quality and cost look like on each, and can describe the switch as work rather than as catastrophe. Bonus if a cheaper model already handles part of the traffic.
Worry if
Confidence without evidence: "we could switch easily". Ask when they last tried. If the answer is never, they do not know, and neither do you.
Can they prove it works?
A demo proves the system can be right once. This section is about whether anyone can tell when it stops being right, which is the difference between a product and a prototype that got funded.
“How do you know a change made the product better rather than worse?”
What it tests. Whether an evaluation set exists. This is the question that most reliably separates teams who have shipped from teams who have demoed.
Sounds like
They describe a set of real inputs with expected outputs, run automatically, with a number they watch. They can tell you what the number is now and what it was three months ago. Strongest answer: they can name a change they reverted because the number went the wrong way.
Worry if
"We test it thoroughly." "Our users tell us." Both mean there is no evaluation set, which means every model upgrade is a gamble and nobody in the company can tell you whether quality is improving or drifting.
“What percentage of the time is it wrong, and how do you know?”
What it tests. Whether they have measured their own failure rate. A founder who has not is not necessarily hiding anything, but they cannot tell you what the support burden looks like at ten times the customers.
Sounds like
A number, a definition of wrong that they can defend, and how it was measured. Even a rough number honestly derived is a strong signal. Better still if they distinguish between wrong and unhelpful, because those are different problems with different fixes.
Worry if
"Very rarely." "It is about as good as a human." The second one is worth pushing on: measured against which humans, doing which task, judged by whom.
“Are the benchmark numbers in the deck yours, or the model provider’s?”
What it tests. Whether the impressive figures describe this company at all. It is a surprisingly common substitution and it is rarely deliberate.
Sounds like
Theirs, measured on their own task, with the method described. They should be able to say why their number differs from the published one for the base model.
Worry if
Published model benchmarks presented as product performance. A model scoring highly on a public benchmark says almost nothing about whether this product answers questions about your customer’s contracts correctly.
Does it make money at scale?
AI features carry a cost per use that traditional software does not. A company can grow revenue and lose more money per customer as it does, and the founders may not know.
“What does one customer cost you in inference per month, and what do they pay?”
What it tests. Whether gross margin is real. This is where AI companies differ structurally from software companies, and where a growing top line can be hiding a widening hole.
Sounds like
Two numbers, promptly, with the heaviest cohort named separately from the average. A founder who runs this weekly will have it to hand.
Worry if
An average that has never been broken down. Averages hide the power users, and in usage-heavy products a small share of customers frequently costs more than they pay.
“What happens to your margin if usage per customer doubles?”
What it tests. Whether the pricing model and the cost model point the same way. Per-seat pricing on a product with per-use costs breaks quietly, at the point where customers start liking it.
Sounds like
They know, and the answer is either that the pricing is usage-linked, or that they have modelled the ceiling and know where it is. Either is fine. Knowing is the test.
Worry if
"More usage means more value, so customers will pay more." That is a hope about a future renegotiation, not a margin.
“Has your cost per unit of work gone up or down over the last six months?”
What it tests. Whether they are actively managing the bill or riding provider price cuts and calling it progress.
Sounds like
Down, with the reasons named: caching, a cheaper model on part of the traffic, shorter prompts, deduplicated requests. They can separate what they did from what the market did.
Worry if
Down, with no explanation. Provider prices have fallen; that is not an achievement of theirs and it will not continue indefinitely.
What happens when it breaks?
Every one of these systems is wrong sometimes. What separates them is whether anybody finds out before the customer does.
“A customer says it gave a wrong answer yesterday. Walk me through what you do.”
What it tests. Whether the system is observable. Ask it as a scenario rather than as a question about tooling, because the scenario is harder to answer from a slide.
Sounds like
A specific sequence: find the request, read the trace, see what was retrieved and what was sent to the model, reproduce it, add it to the evaluation set. The last step is the one that distinguishes a mature team.
Worry if
Hesitation, or an answer that ends at "we would look into it". If they cannot reconstruct a single interaction, they cannot debug the system, they can only adjust the prompt and hope.
“What can it do that you would not want it to do, and what stops that?”
What it tests. Whether anybody has thought about blast radius. It matters most where the system takes actions rather than only producing text.
Sounds like
A named list of the dangerous actions, with the specific controls: approval steps, spend limits, reversal paths, a hard stop. They have thought about the worst case rather than the average case.
Worry if
"The model does not do that." Models do do that, occasionally, and the question is what happens when it does rather than how unlikely it is.
“What breaks first at ten times the current traffic?”
What it tests. Whether they have a real mental model of their own system. The answer matters less than whether one exists.
Sounds like
A specific component, a reason, and a rough idea of the work to fix it. Frequently it is a rate limit at a provider or a retrieval index that stops being fast, and knowing which is the point.
Worry if
"It scales." Nothing scales. Everything has a next bottleneck, and a founder who cannot name theirs has not looked.
“Where does customer data go, and does the model provider train on it?”
What it tests. A contractual and reputational risk that founders sometimes answer wrongly in good faith, because they read the marketing page rather than the terms.
Sounds like
A clear description of the boundary, which providers are involved, what is retained and for how long, and the specific contractual position on training. They can point at the clause.
Worry if
A confident "no, it is private" with no detail. Ask which provider, on which plan, because the answer differs between them and between tiers of the same one.
Who actually built it?
The most expensive risk in an early technical company is concentration, and it is the easiest one to miss because the person carrying it is usually in the room being impressive.
“Which one person leaving would hurt most, and what would break?”
What it tests. Concentration risk. Ask it directly. Founders usually answer this one honestly because the person is on their mind already.
Sounds like
A name, a specific area, and something being done about it: pairing, documentation, a hire in progress. Concentration in an early company is normal. Being unaware of it is not.
Worry if
"The team is cross-functional, nobody is a single point of failure." In a company of this size that is almost never true, and the answer suggests either evasion or that nobody has looked.
“Has anyone here run an AI system in production before this one?”
What it tests. Whether the hard lessons have been paid for already or are still ahead. It is not disqualifying either way. It changes the timeline.
Sounds like
Yes, with specifics about what went wrong last time and what they do differently now. The war story is the signal.
Worry if
Extensive experience with models in a research or a demo context presented as production experience. The gap between the two is monitoring, cost control and failure handling, and it is usually a quarter of work nobody planned for.
“Which parts were built by people who no longer work here?”
What it tests. Whether the codebase has orphaned regions, which is common after an agency build or an early contractor and is rarely volunteered.
Sounds like
A straight answer identifying which parts, and confirmation that someone current owns and understands them now.
Worry if
Surprise at the question, or a vague answer. Orphaned code is not fatal and it is a real cost, and it should be in the technical plan rather than discovered in month two.
Where the risk sits
Nothing asked yet in
- Is any of this theirs?
- Can they prove it works?
- Does it make money at scale?
- What happens when it breaks?
- Who actually built it?
What this does not do
This sorts questions into answered well, answered badly, and not asked. It deliberately does not produce a score or a recommendation. 16 answers cannot be turned into an investment decision, and any tool that claims otherwise is selling you certainty it does not have. What it gives you is a map of where the remaining technical risk sits, in a form you can put in a memo.
If you would rather not ask
The first memo is free, one per fund, inside 48 hours.
A questionnaire is only as good as the person using it. Where the answers need checking against the actual codebase rather than taken on trust, that is the thing we do.
Technical due diligence memo
Whether the codebase matches the story, where the team carries risk it has not priced, and what the first six months of engineering will cost. It states its own gaps on the first page rather than implying coverage it does not have.
One limit worth stating
Every question here can be answered well by a company that is still going to fail, and answered badly by one that is going to work. A team with no evaluation set and a product customers cannot stop using is a better investment than a rigorous team with no customers, and nothing on this page will tell you that. This covers technical risk only, which is one input to a decision made mostly on other things.
The specific thing it cannot do is verify an answer. A founder can describe an evaluation set convincingly without having one. That is why several follow-ups ask to be shown a thing rather than told about it, and it is the difference between a questionnaire and reading the codebase.
FAQ
Questions about this list
Next step
If a company you are looking at answered three of these badly and you want somebody to check whether it matters, that is what the free memo is for.