RC
ArchitectureArtificial IntelligenceSoftware ArchitectureTechnology

Has AI Flipped the Build vs Buy Decision?

An engineer with a good agentic setup can now stand up in a fortnight what used to take a quarter. The demo is real. But "we can build it faster" and "we should build it" are different claims, and the gap is where a lot of money is currently being set on fire. This post works through the five-year TCO arithmetic, a nine-gate decision tree, and the new failure mode nobody has priced yet: owning a system nobody had time to understand.

Harish Kumar
Share

Has AI Flipped the Build vs Buy Decision?

Short answer: no. It moved the line, it didn't erase it — and the part it moved is the part that was never expensive.


Every vendor renewal I've been in this year eventually arrives at the same sentence, usually delivered by someone on my own side of the table: "Couldn't we just build this ourselves now?"

It's a fair question. An engineer with a good agentic coding setup can stand up in a fortnight what used to be a quarter of work. The demo is real. The velocity is real.

But "we can build it faster" and "we should build it" are different claims, and the gap between them is where a lot of money is currently being set on fire.

This post is my attempt to write down the decision framework properly — with the arithmetic, the failure modes, and the places where the answer genuinely has flipped.


1. The mistake: optimising the smallest term

The build-vs-buy conversation has always over-indexed on the build, because the build is the part with a visible price tag. It has a start date, a team, a burn rate, a Gantt chart. It's legible.

The rest of the cost is diffuse. It shows up as someone's Tuesday afternoon, a quarterly audit, a CVE that lands on a Friday, a dependency upgrade that eats a sprint, a resignation letter from the one person who understood the batch job.

Here's the shape of a five-year total cost of ownership for a non-trivial platform capability in a regulated environment. Numbers are illustrative — substitute your own — but the proportions are consistent with what I've seen:

pie showData
    title Five-year TCO of a built capability (illustrative)
    "Initial build" : 25
    "Run and operate (on-call, incidents, capacity)" : 22
    "Maintenance (upgrades, CVEs, dependency drift)" : 18
    "Integration and change (upstream and downstream)" : 20
    "Compliance, audit, evidence, access reviews" : 15

The initial build is a quarter. Everything else is the tail.

Now apply AI to it. Be generous — assume a 45% reduction in build effort, which is at the optimistic end of what I've measured on real work rather than on greenfield demos.

graph LR
    subgraph Before["Without AI — 100 units"]
        B1["Build<br/>25"]
        B2["Run + Maintain + Integrate + Comply<br/>75"]
    end
    subgraph After["With AI — 89 units"]
        A1["Build<br/>14"]
        A2["Run + Maintain + Integrate + Comply<br/>75"]
    end
    Before --> R["Net TCO change<br/><b>-11%</b>"]
    After --> R

    style B1 fill:#1B2A32,color:#fff
    style A1 fill:#1B2A32,color:#fff
    style B2 fill:#9DACA7,color:#000
    style A2 fill:#9DACA7,color:#000
    style R fill:#C13629,color:#fff

A 45% cut to the build produces an 11% cut to TCO.

Eleven percent is real money and I'll take it. It is not a decision flip. If an 11% TCO delta reverses your build-vs-buy conclusion, the two options were near-identical to begin with and you should be deciding on strategic grounds, not spreadsheet grounds.

Why the tail doesn't compress the way the build does

This is the crux, so it's worth being specific about why AI doesn't dent the other 75%:

Cost driver Why AI barely touches it
On-call and incident response The cost is human availability at 2am and accountable decision-making under pressure, not the cost of typing a fix.
Compliance and audit evidence Auditors need attested human accountability and a control owner. An AI-generated control narrative still needs a named person to sign it.
Integration churn Your cost is driven by other teams' release cadences and their undocumented behaviour. You don't control the input rate.
Dependency and platform upgrades The work is regression risk and coordination across consumers, not code generation. AI helps at the margins; the change window is the bottleneck.
Key-person risk AI-generated code that nobody on the team deeply understands makes this worse, not better. See section 4.
Product decisions Roadmap, prioritisation, and saying no. Entirely unaffected.

Meanwhile — and this gets forgotten — your vendor has the same AI. Their build got cheaper too. Some of that shows up as a faster roadmap, some as margin, some as competitive price pressure when you push. The comparison moves on both sides of the equation.


2. What genuinely has changed

If I only argued "nothing has changed" I'd be wrong, and dismissively so. Three things really did move.

2.1 The floor dropped

There has always been a band of capability that was too small to justify a project but too annoying to live without. An internal admin portal. A reconciliation script with a UI. A self-service request form that saves a team forty tickets a month. A one-off migration tool.

Historically these lost to "just buy a SaaS seat" or, more often, to "just live with the spreadsheet," because the fixed overhead of starting a project — the design doc, the review, the pipeline, the onboarding — swamped the value.

That overhead collapsed. Below that line, build genuinely won.

graph TD
    A["Capability under consideration"] --> B{"Effort to build,<br/>post-AI"}
    B -->|"Days"| C["<b>BUILD</b><br/>Below the old floor.<br/>Not worth a procurement cycle."]
    B -->|"Weeks to months"| D{"Does it touch regulated data,<br/>money movement, or customers?"}
    D -->|No| E["<b>Lean build</b><br/>Low tail cost. Accept the ownership."]
    D -->|Yes| F["Full TCO analysis applies.<br/>Continue to Section 3."]

    style C fill:#1B2A32,color:#fff
    style E fill:#4A6670,color:#fff
    style F fill:#C13629,color:#fff

The caveat, which nobody puts on the slide: the floor dropping means you will accumulate far more small systems than before. Fifty afternoon-projects is not fifty afternoons of cost. It's fifty things to patch, fifty things to inventory, fifty things to hand over. The floor dropping is a real win and the most likely source of the next decade's technical debt. Set a sunset policy on day one or don't build it.

2.2 Buy-side leverage improved

This is the most underrated effect and the one I'd actually monetise.

You can now credibly say "we've prototyped this internally" — and mean it — during a renewal. That changes the conversation materially. Vendors price against your alternatives, and your alternative just got cheaper and, crucially, demonstrable.

My advice: spend that leverage in the negotiation, not in the architecture. A working prototype that wins you 30% off a five-year contract is worth more than the same prototype promoted to production and owned forever. Build to negotiate, then throw it away without sentiment.

2.3 The differentiation test got sharper

When building was expensive, cost was doing the filtering for you. It stopped bad build decisions automatically. That filter is now much weaker, which means the strategic test has to do all the work.

The test is unchanged, it's just load-bearing now: would a customer ever notice this is yours?

If the honest answer is no, an AI-accelerated build just gets you a bespoke system, owned forever, that does what a commodity product does — but with your logo on the maintenance bill.


3. The decision framework

Here's how I actually run it now. The gates are ordered deliberately: the cheap disqualifiers come first, and cost analysis comes last, because cost is the argument people reach for when they've already decided emotionally.

flowchart TD
    START["Capability request"] --> G1{"Is this a<br/>differentiator?<br/>Would a customer notice?"}

    G1 -->|"Yes — it's our edge"| G2{"Does a mature<br/>product already do<br/>80% of it?"}
    G1 -->|"No — commodity"| G6{"Does a viable<br/>product exist?"}

    G2 -->|No| BUILD1["<b>BUILD</b><br/>Core differentiator,<br/>no market answer"]
    G2 -->|Yes| G3{"Can we extend it<br/>via supported APIs<br/>without forking?"}

    G3 -->|Yes| HYBRID["<b>BUY + EXTEND</b><br/>Buy the platform,<br/>build the thin<br/>differentiating layer"]
    G3 -->|No| BUILD2["<b>BUILD</b><br/>Extension ceiling<br/>caps our upside"]

    G6 -->|No| G7{"Effort to build,<br/>post-AI?"}
    G6 -->|Yes| G8{"Regulatory,<br/>data residency, or<br/>exit-risk blocker?"}

    G7 -->|"Days"| BUILD3["<b>BUILD</b><br/>Below the floor"]
    G7 -->|"Months"| G9{"Can we defer<br/>or descope<br/>entirely?"}

    G9 -->|Yes| DEFER["<b>DEFER</b><br/>The cheapest system<br/>is the one not built"]
    G9 -->|No| BUILD4["<b>BUILD — minimum viable</b><br/>with a named owner<br/>and a sunset date"]

    G8 -->|Yes| G10{"Is the blocker<br/>negotiable or<br/>mitigable?"}
    G8 -->|No| G11{"Five-year TCO<br/>comparison"}

    G10 -->|Yes| G11
    G10 -->|No| BUILD5["<b>BUILD</b><br/>Constraint-driven,<br/>eyes open on cost"]

    G11 -->|"Buy within 1.3x of build"| BUY["<b>BUY</b><br/>Transfer the tail.<br/>Use AI prototype as<br/>negotiating leverage"]
    G11 -->|"Buy above 1.3x of build"| G12{"Is the delta driven by<br/>seat count that will<br/>keep growing?"}

    G12 -->|Yes| BUILD6["<b>BUILD</b><br/>Per-seat economics<br/>break at our scale"]
    G12 -->|No| BUY2["<b>BUY</b><br/>Premium is smaller<br/>than the ownership tail"]

    style BUILD1 fill:#1B2A32,color:#fff
    style BUILD2 fill:#1B2A32,color:#fff
    style BUILD3 fill:#1B2A32,color:#fff
    style BUILD4 fill:#1B2A32,color:#fff
    style BUILD5 fill:#1B2A32,color:#fff
    style BUILD6 fill:#1B2A32,color:#fff
    style BUY fill:#4A6670,color:#fff
    style BUY2 fill:#4A6670,color:#fff
    style HYBRID fill:#6B8E7F,color:#fff
    style DEFER fill:#C13629,color:#fff

Three notes on the gates that tend to generate argument:

The 1.3x threshold. Why not 1.0x? Because a build TCO estimate is systematically optimistic and a buy quote is contractually binding. The 30% band is a humility premium on your own forecast. If your organisation has a track record of accurate delivery estimates, tighten it. Mine doesn't, so I don't.

"Can we defer entirely?" This gate closes more requests than any other, and it belongs in the flow explicitly. AI making something cheap to build is not a reason to build it. Every system you don't create has a TCO of zero, forever.

Buy + extend is the fastest-growing answer. Post-AI, the thin differentiating layer on top of a bought platform is dramatically cheaper than it used to be. This is where AI has changed my decisions most in practice — not by flipping buy to build, but by making the hybrid viable where previously you took the vendor's workflow as given.


4. The new failure mode: velocity without comprehension

There's a risk that didn't exist at this scale before, and it deserves its own section because it's the one that will hurt people.

When a build takes six months, understanding accretes as a by-product. Design debates, code review arguments, and the sheer time spent staring at the thing produce a team that genuinely knows how it works. That understanding is the asset that makes year three survivable.

When the build takes three weeks, that by-product doesn't form. You end up owning a system that works, that nobody has argued about, and that no one can confidently reason about during an incident.

graph LR
    A["Fast AI-assisted build"] --> B["System in production"]
    B --> C["Comprehension debt<br/>Nobody deeply owns<br/>the design rationale"]
    C --> D["Incident at 02:00"]
    D --> E["Mean time to<br/>resolution climbs"]
    C --> F["Change requests<br/>take longer than<br/>the original build"]
    F --> G["Effective TCO<br/>exceeds the buy option"]
    E --> G

    style A fill:#4A6670,color:#fff
    style C fill:#C13629,color:#fff
    style G fill:#C13629,color:#fff

Mitigations that have worked for me:

  • Design review before generation, not after. The architecture decision record is written and argued by humans. The implementation can be accelerated. Reversing that order is where the debt comes from.
  • A named owner per system, recorded, with the on-call rota as the enforcement mechanism. If nobody will carry the pager for it, it doesn't ship.
  • Runbooks written by the person who will use them at 2am, tested in a game day. AI can draft; the drill is what validates.
  • A sunset date at creation. Every fast build gets an expiry review. Most should die. The ones that survive get promoted deliberately, with proper ownership and a proper budget.

5. Where I've landed

The rule I've used for years hasn't changed much: build what customers would notice, buy what they'd never see.

What AI changed:

  • The floor moved down. More small things are worth building. Manage the resulting sprawl deliberately or it becomes your next debt programme.
  • The hybrid got better. Buy the platform, build the thin differentiating edge. This is where most of the genuine gain is.
  • Your negotiating position improved. Spend that in procurement, not in the architecture.
  • The cost filter weakened, so the strategic filter has to be stronger and applied more honestly.

What it didn't change: you own everything you build. AI writes the code. It doesn't sign the audit, take the 2am call, absorb the regulatory accountability, or answer to the board when the thing is down during a claims surge.

The build was never the expensive part. It's just the part with a price tag.


What's your experience? I'm particularly interested in cases where the flip was genuine — where post-AI TCO actually reversed a buy decision that held pre-AI. I suspect they exist and are concentrated in high-seat-count SaaS. I'd like to see the numbers.

Related reading