How to assess a startup codebase in two days
What technical due diligence can establish in 48 hours, what it cannot, and the six questions that separate a platform from an API wrapper.
A partner has four days before the committee meets and the deck says the platform is proprietary. Somebody has to open the repository and find out what that word is doing.
Two days is enough to answer several questions that change a decision, and not enough to answer several others. Knowing which is which is most of the skill, and a memo that fails to separate them is worse than no memo, because it gives a committee confidence it has not earned.
What two days can establish
Whether the technical claim is supported. If the deck says proprietary models, two days is enough to see whether there are training pipelines, evaluation sets and model artefacts, or whether there is an API client with a well-written prompt. Both can be good businesses. They carry different defensibility and they should carry different valuations.
Where the knowledge is concentrated. Commit history tells you how many people have meaningfully touched the core system. Where one engineer authored eighty per cent of it and there is no documentation, you have priced a key-person risk whether or not you meant to.
Whether the infrastructure cost scales with users or worse. Read the architecture for per-request model calls, unbounded retries, and vector searches over collections that grow with the customer base. This is arithmetic, it takes an afternoon, and it regularly changes a unit economics story.
What the first six months of engineering cost. Headcount to hold the current system, plus headcount to build the roadmap in the deck, against the growth assumptions the company has already given you. Founders are usually optimistic here by a predictable factor.
Whether the team can ship. Deployment frequency, whether tests exist and run, whether pull requests are reviewed. A team that deploys weekly with a passing CI pipeline is a different asset from one that deploys when the founder is awake.
What two days cannot establish
Say these out loud in the memo, on the first page.
Security. Two days finds obvious exposures, a secret in the repository, an open bucket, an unauthenticated endpoint. It does not constitute a security audit, and reporting the absence of obvious problems as a clean bill is the single most dangerous sentence such a memo can contain.
A codebase over roughly 200,000 lines. Beyond that, two days produces an impression of a subset. The memo should say which subset was read.
Model quality that rests on proprietary data. Where the claim is that the model is better because the training data is better, and the data cannot be inspected, no amount of code reading resolves it. That question goes to customer references instead.
Whether the architecture is right. Architecture is judged against a roadmap and a scale, both of which are claims by the company. The most a memo can honestly say is whether the architecture is consistent with the roadmap the company has described.
The six questions
These are for the founder, phrased so the answers are checkable afterwards.
- Which parts of the system would you rewrite if you had a free quarter, and why? Founders answer this honestly more often than expected, and the answer names the real technical debt faster than reading does.
- If your primary model provider doubled its price tomorrow, what changes? Tests whether provider dependency has been thought about, and how tightly the product is coupled to one vendor's behaviour.
- Who is the second person who understands the core system, and what happens in the fortnight after the first one leaves? Concentration risk, asked in a way that is hard to deflect.
- What is your inference cost per active user this month, and what was it three months ago? The trend matters more than the level. A cost per user rising with scale is a different company from one falling.
- Show me the last production incident and what changed afterwards. Reveals whether there is an incident process at all, which correlates with almost everything else.
- What does your evaluation set look like, and when did it last catch a regression? A team that cannot answer has no way of knowing whether a prompt change made the product worse, which means every model update is a coin flip.
Writing it up
Three to six pages, and separate what was verified from what was inferred. A memo that reads as uniformly confident is less useful than one that marks its own uncertainty, because a partner needs to know how much weight each claim carries.
State the access you had. "Read access to the main repository, no production access, no conversation with the engineering team" tells the committee what the memo is worth in one line.
Where you found nothing concerning, say that you found nothing concerning within the scope you covered, and restate the scope. The difference between "no problems" and "no problems in the two systems I read" is the difference between a memo that ages well and one that gets quoted back at you.
The uncomfortable part
Diligence written in two days is sometimes wrong, and the failure mode is asymmetric. A false alarm costs a founder an awkward meeting. A missed issue costs a fund a mark-down and costs you the relationship.
Given that asymmetry, the correct bias is toward flagging with uncertainty stated rather than staying silent to appear decisive. A memo listing three concerns marked as unverified, with the checks that would settle each, is more useful to a committee than one confident conclusion. It is also less impressive to read, which is why the incentive runs the other way.
If you are running a process this month, the fastest way to see whether this shape of memo helps is to send a live one. The 48-hour memo is free and states its own limits, and the full description of what it covers is on that page.
If you are evaluating a company and want a technical read before the next partner meeting, that is the conversation to have.
Request the memoWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- Best AI consultancies for startups, 2026Seven AI partners a startup can actually hire, what each one is genuinely for, and the two questions none of them answer on their websites.
- Best AI training providers for teams, 2026Six ways to train a team on AI, what each one actually teaches on, and why the productivity figures in this market are so hard to compare.
- Is Omarchy the AI-native OS for power users?Omarchy puts coding agents in the operating system rather than in a browser tab. That is the right direction, and it makes one missing piece obvious.
Work with us
Looking at a company right now and want a technical read before the partner meeting?
Request the free memoGet the next one
New writing in Technical due diligence as it goes up, roughly twice a month. One article per email and nothing else in it.