The most common reason an AI platform business case fails CFO review is not a bad model—it is the wrong baseline. Measuring platform value against squads doing nothing, rather than the best realistic local alternative, inflates the ROI by design. This post derives incremental platform value across five benefit terms, presents three-scenario phased cash flows over a 3-year horizon, and resolves the kill rule vs. patient capital tension through real-options logic.

Underwriting the Return on an AI Platform

Part 3 of 3 — AI Platform Budget Decisions


A business case that hides assumption risk inside the model architecture is not a financial projection. It is a narrative with a spreadsheet attached.


Why Most Platform Business Cases Fail CFO Review

The flawed pattern is familiar:

20 Engineers × 15% Saved × $200,000 Loaded Cost = $600,000 Annual Return

A CFO rejects this immediately—and correctly—for two reasons.

First, saved engineering hours do not reduce payroll. Redeployable capacity creates opportunity value but not cash savings unless headcount or contract spend is actually reduced.

Second, the formula omits coordination overhead entirely. A centralized platform introduces alignment meetings, intake queues, and contract negotiation cycles. These costs erode theoretical savings and rarely appear in the business case.

The correction that matters most, though, is not to the formula. It is to the baseline.


What Is the Right Baseline?

Incremental Platform Value = Outcomes(federated platform) − Outcomes(best feasible local architecture)

The best feasible local alternative is not squads building everything poorly from scratch. It is squads buying managed model access, purchasing commodity gateway and tracing tools, retaining local capability contracts, and independently handling governance.

Under that baseline, duplicate plumbing costs vanish, and baseline telemetry is already handled. The incremental value of the shared platform is narrower than the full ROI column suggests. It is also the only number a PM can defensibly advocate.

The Baseline Moves

The counterfactual is not static. The commercial market for LLM developer infrastructure is improving aggressively. Capabilities that required custom squad engineering twelve months ago—structured JSON enforcement, automated prompt caching, basic evaluation harnesses—are now native features of managed model APIs.

If the platform's value rests on providing capabilities that vendor APIs will commoditize next quarter, the platform's incremental value shrinks over time. A defensible platform business case must derive value from three durable sources:

Building internal software that merely wraps model calls is capital spent fighting the vendor roadmap.


The Five Benefit Terms—Derived, Not Asserted

A CFO will not accept asserted benefit figures. A label of "High Confidence" does not substitute for a visible derivation chain. Below is the explicit methodology for all five benefit terms, grounded in a reference organization: 4 squads, 20 engineers at $200k fully loaded cost per year, processing 1.2 million multi-page documents and 50 billion tokens annually.

1. Avoided Rework

Product squads maintaining local prompt wrappers, custom retry logic, and independent evaluation scripts spend real engineering capacity on non-differentiating plumbing.

$$C_{\text{rework}} = N_{\text{squads}} \times \text{FTE}_{\text{plumbing}} \times \text{Loaded Cost} \times \text{Avoidance Factor} \times \text{Cash Realization}$$

2. Risk Reduction

Centralization standardizes safety filters and evaluation assertions, reducing localized failure rates—but it introduces shared failure risk. The formula nets both.

$$\Delta R_{\text{risk}} = E[L_{\text{local baseline}}] - E[L_{\text{shared platform}}] - E[L_{\text{correlated failure}}]$$

In the base case: each squad has a 20% annual probability of a formatting failure or hallucination reaching customers at $150k remediation cost. The platform reduces that to 3%. But a misconfigured gateway rule now impacts all four squads simultaneously at 4% annual probability and $450k systemic cost.

The correlated failure term is not optional arithmetic. A platform that improves per-squad incident rates while introducing a common-mode failure vector has not reduced risk—it has redistributed it.

3. Compliance Exposure

Compliance exposure isolates statutory regulatory penalties and external audit findings, separate from the operational remediation costs modeled in risk reduction.

$$R_{\text{compliance}} = P(\text{statutory breach}) \times \text{Expected Statutory Fine} \times \text{Mitigation Efficacy}$$

4. Distillation Savings (Net)

Inference cost reduction achieved by replacing generalist frontier models with task-distilled local models.

For the reference workload (50 billion annual tokens; 60% eligible for distillation; $2.35 delta cost per million tokens):

Gross Savings = 50,000M tokens × 0.60 × $2.35/M = $70,500/year

Subtracting operational costs (fine-tuning, hosting, evaluation, maintenance, fallback)—which equal ~50% of gross savings in Year 1:

5. Churn Protection


The Scenario Model

Benefit Term Conservative Base Upside Evidence Quality Confidence Condition
Avoided rework $150k $320k $450k High Audited contractor spend reduction or signed backlog reallocations
Risk reduction $20k $84k $134k Medium Net of correlated-failure term; requires prior error incident baseline
Compliance exposure $0 $50k $100k Medium External statutory fines only—no overlap with operational risk
Distillation savings $0 $35k $70k Low–Medium Net of fine-tuning, hosting, and eval costs; zero in Year 1 conservative
Churn protection $0 $0 $65k Low Upside only; requires documented customer contract SLA link
Total Annual Gross Benefit $170k $489k $819k

Three Years of Phased Cash Flows

A flat Year 1 table obscures the true capital profile. Upfront build costs are front-loaded; distillation, learning, and compounding benefits materialize in later years. Below is the multi-year schedule evaluated at a 10% WACC, with a $140k initial build cost and $54k Year 1 / $40k Year 2–3 operating overhead.

Conservative

Cash Flow Line Year 1 Year 2 Year 3
Gross Realized Benefits $170,000 $190,000 $210,000
Platform Cost (Cap + Op) ($194,000) ($40,000) ($40,000)
Net Annual Cash Flow ($24,000) $150,000 $170,000
Discounted Cash Flow ($21,818) $123,967 $127,724
3-Year NPV @ 10% WACC $229,873

Base Case

Cash Flow Line Year 1 Year 2 Year 3
Gross Realized Benefits $489,000 $540,000 $600,000
Platform Cost (Cap + Op) ($194,000) ($40,000) ($40,000)
Net Annual Cash Flow $295,000 $500,000 $560,000
Discounted Cash Flow $268,182 $413,223 $420,736
3-Year NPV @ 10% WACC $1,102,141

Upside Case

Cash Flow Line Year 1 Year 2 Year 3
Gross Realized Benefits $819,000 $920,000 $1,050,000
Platform Cost (Cap + Op) ($194,000) ($40,000) ($40,000)
Net Annual Cash Flow $625,000 $880,000 $1,010,000
Discounted Cash Flow $568,182 $727,273 $758,828
3-Year NPV @ 10% WACC $2,054,283

Two things the phased schedule reveals that the flat table hid:


Resolving the Kill Rule vs. Patient Capital Tension

A fundamental tension runs through every AI platform investment:

If conservative returns are negative in Year 1 (−$24k) because distillation and learning benefits require multiple cycles to compound, a rigid two-cycle kill rule will execute the platform precisely when it needs patient capital.

Platform Stages as Call Options

The resolution is to replace point-estimate capital budgeting with real-options valuation.

flowchart TD
    INVEST["**Stage 1 / Stage 2 Investment**\nPurchases Call Option on Shared Leverage"]
    WINDOW["90-Day Evidence & Option Window"]
    INVEST --> WINDOW
    WINDOW --> ALIVE
    WINDOW --> KILL

    ALIVE["**Uncertainty Unresolved**\n• Schema overlap promising\n• Active squad feedback\n• Low marginal pod cost"]
    KILL["**Hypothesis Invalidated**\n• Squads bypass platform\n• Schemas structurally diverge\n• Overlays exceed shared core"]

    ALIVE --> EXTEND["**Keep Option Alive**\nExtend Stage 2 Pod\nDo NOT advance to Stage 3"]
    KILL --> EXIT["**Exercise Kill**\nDecommission contract\nAbsorb reversal cost early"]

    style INVEST fill:#1e293b,stroke:#60a5fa,stroke-width:2px,color:#f8fafc
    style WINDOW fill:#1e293b,stroke:#94a3b8,stroke-width:1px,color:#cbd5e1
    style ALIVE fill:#064e3b,stroke:#10b981,stroke-width:1.5px,color:#f8fafc
    style KILL fill:#450a0a,stroke:#f87171,stroke-width:1.5px,color:#f8fafc
    style EXTEND fill:#065f46,stroke:#34d399,stroke-width:1px,color:#f8fafc
    style EXIT fill:#7f1d1d,stroke:#fca5a5,stroke-width:1px,color:#f8fafc

Stage 1 and Stage 2 are call options, not full infrastructure commitments. They purchase the right—but not the obligation—to expand to a Stage 3 shared service once uncertainty resolves. The value of the pilot is information value: confirming whether squads share evidence semantics, verifying operator override volume, and testing squad collaboration.

The options-aware kill and extension rule:


Where This Approach Fails

A decision framework that does not identify its own failure conditions is marketing copy.

1. Workload Scarcity. One or two AI features with low transaction volume cannot amortize platform build and coordination costs. The conservative scenario produces negative net returns. Do not build a platform. Revisit when volume and semantic overlap mature.

2. The Broken Learning Flywheel. Distillation depends on an operational chain: production operation → operator override → validated correction → curated evaluation corpus → distilled model. If operators correct inconsistently, enterprise customers restrict cross-tenant data pooling, or sparse overrides fail to reach statistical significance, the chain breaks. Gate distillation savings to the upside scenario only until automated telemetry verifies closed-loop curation in production.

3. The Moving Model Frontier. Frontier model capabilities improve materially every six to twelve months. A platform that encodes custom prompt patterns, bespoke routing heuristics, or fragile retrieval wrappers finds that an upstream model release renders eighteen months of internal engineering obsolete. Decouple platform contracts from model implementation details. Anchor leverage in proprietary organizational state and task-specific evaluation suites.

4. The False Reuse Trap at Scale. Centralization gates are relaxed under organizational pressure, grouping superficially similar workflows into a shared abstraction. The platform team spends its roadmap maintaining dozens of configuration flags and divergent assertion suites for incompatible squads. Enforce the 2-Week Convergence Spike and continuously track the reuse ratio. Decommission shared contracts when domain overlays exceed common logic.

5. Governance Gridlock. Centralization concentrates institutional risk, attracting multi-stakeholder oversight. Every squad release requires sequential approvals from model risk governance, data privacy, security architecture, and the platform team. Technical fragmentation is replaced with organizational paralysis. Automate governance controls directly at the platform gateway layer, tier controls by risk classification, and preserve autonomous self-service for low-risk workflow iterations.

6. The Compliance Blanket Failure. Treating regulatory compliance as an unchallengeable blank check bypasses financial scrutiny. CFOs audit compliance-driven platform expenditures and demand verifiable proof of control efficacy. Underwrite compliance defensively using actuarial expected loss ($P(\text{Breach}) \times \text{Fine} \times \text{Efficacy}$) rather than qualitative assertions.


How Product Managers Know the Platform Is Working

A platform that cannot be measured cannot be governed.

Metric Category Metric Target Threshold
Adoption Workflow adoption rate ≥ 75% of eligible squads by month 6
Adoption Platform bypass rate < 5%
Efficiency Time to first capability invocation < 3 business days from squad onboarding
Efficiency Contract-change lead time < 5 business days
Quality Evaluation coverage 100% of production inference volume
Quality Escaped regression rate < 1% of deployments
Economics Realized unit cost Downward trend Q/Q
Learning Correction retention rate ≥ 40% of overrides promoted to eval suites
Learning Reuse ratio < 0.25 (above signals Reuse Trap)
Governance Domain-pod autonomy score ≥ 90% of pod releases without central approval
Reversibility Decoupling lead time ≤ 10 engineer-days per consuming squad

Watch for four tripwires that signal the platform has expanded beyond its demonstrated reuse boundary:


The Decision Gate: Three Outcomes

Outcome Conditions Action
Proceed All required centralization gates pass. Three or more scored dimensions demonstrate meaningful overlap. Base-case scenario produces positive 3-year NPV against the best local alternative. Advance to next platform stage. Assign capability steward. Establish operating metrics and option review milestones.
Pilot Required gates pass but scored overlap is borderline, or the base-case ROI requires validating assumptions through a narrow staged release. Advance to a time-limited Stage 2 pilot with two squads. Fund the 2-week Convergence Spike. Price as a call option with a 90-day review gate.
Remain Local A required gate fails, scored overlap is structurally divergent, or the conservative scenario does not justify investment without an unvalidated compliance excuse. Retain capability inside the product squad. Continue managed model access. Re-evaluate when volume, reuse, or risk classification warrants a new assessment.

One-Line Synthesis

The platform must earn the right to expand—from managed access, to domain capability, to shared service—through demonstrated reuse, measurable outcomes, and realized economics. Sequence capital on evidence, price early stages as real options, and never confuse architectural enthusiasm with financial return.



The ideas in this post are my own — they emerged from questions I asked while learning applied AI concepts and putting them to work in my job and my projects. The prose was developed with AI assistance.

Frequently Asked Questions

What is the correct baseline for an AI platform ROI calculation?

The baseline for platform ROI is not 'squads doing nothing' or 'squads building everything from scratch poorly.' The correct baseline is the best realistic local alternative: product squads buying managed model access, retaining local capability contracts, and independently handling governance. Platform incremental value is the difference in cost, risk, and learning outcomes between the federated platform model and that best local alternative.

How do you resolve the tension between the Kill Rule and patient capital in AI platforms?

Frame staged platform investment as purchasing a real option on enterprise leverage rather than a point-estimate NPV commitment. Early stages purchase information value—reducing uncertainty about cross-squad semantic overlap and adoption friction—before committing to heavy shared service infrastructure. Replace blunt time-based kill gates with an options-aware rule: if uncertainty remains unresolved but squads actively collaborate, extend the low-cost domain pod stage; if the hypothesis is invalidated through contract bypass or structural schema divergence, exercise the kill.

Download the Architecture of Proof Checklist

Ready to implement? Get the definitive checklist for building verifiable AI systems.

Zoomed image
Free Download

Downloading Resource

Enter your email to get instant access. No spam — only occasional updates from Architecture of Proof.

Success

Link Sent

Great! We've sent the download link to your email. Please check your inbox.