Underwriting the Return on an AI Platform
Part 3 of 3 — AI Platform Budget Decisions
A business case that hides assumption risk inside the model architecture is not a financial projection. It is a narrative with a spreadsheet attached.
Why Most Platform Business Cases Fail CFO Review
The flawed pattern is familiar:
20 Engineers × 15% Saved × $200,000 Loaded Cost = $600,000 Annual Return
A CFO rejects this immediately—and correctly—for two reasons.
First, saved engineering hours do not reduce payroll. Redeployable capacity creates opportunity value but not cash savings unless headcount or contract spend is actually reduced.
Second, the formula omits coordination overhead entirely. A centralized platform introduces alignment meetings, intake queues, and contract negotiation cycles. These costs erode theoretical savings and rarely appear in the business case.
The correction that matters most, though, is not to the formula. It is to the baseline.
What Is the Right Baseline?
Incremental Platform Value = Outcomes(federated platform) − Outcomes(best feasible local architecture)
The best feasible local alternative is not squads building everything poorly from scratch. It is squads buying managed model access, purchasing commodity gateway and tracing tools, retaining local capability contracts, and independently handling governance.
Under that baseline, duplicate plumbing costs vanish, and baseline telemetry is already handled. The incremental value of the shared platform is narrower than the full ROI column suggests. It is also the only number a PM can defensibly advocate.
The Baseline Moves
The counterfactual is not static. The commercial market for LLM developer infrastructure is improving aggressively. Capabilities that required custom squad engineering twelve months ago—structured JSON enforcement, automated prompt caching, basic evaluation harnesses—are now native features of managed model APIs.
If the platform's value rests on providing capabilities that vendor APIs will commoditize next quarter, the platform's incremental value shrinks over time. A defensible platform business case must derive value from three durable sources:
- Proprietary institutional data: Verified human override logs, domain corrections, and curated organizational memory.
- Specialized domain evaluation: Task-specific assertion suites and gold-standard corpora that commercial models cannot assess natively.
- Workflow authority boundaries: Deep integration with enterprise compliance, audit trails, and financial execution limits.
Building internal software that merely wraps model calls is capital spent fighting the vendor roadmap.
The Five Benefit Terms—Derived, Not Asserted
A CFO will not accept asserted benefit figures. A label of "High Confidence" does not substitute for a visible derivation chain. Below is the explicit methodology for all five benefit terms, grounded in a reference organization: 4 squads, 20 engineers at $200k fully loaded cost per year, processing 1.2 million multi-page documents and 50 billion tokens annually.
1. Avoided Rework
Product squads maintaining local prompt wrappers, custom retry logic, and independent evaluation scripts spend real engineering capacity on non-differentiating plumbing.
$$C_{\text{rework}} = N_{\text{squads}} \times \text{FTE}_{\text{plumbing}} \times \text{Loaded Cost} \times \text{Avoidance Factor} \times \text{Cash Realization}$$
- Conservative ($150k): 0.25 FTE per squad × 4 squads = 1.0 FTE ($200k loaded). 75% avoidance efficiency; 100% realized as reduced contractor augmentation.
- Base ($320k): 0.50 FTE per squad × 4 squads = 2.0 FTE ($400k loaded). 80% elimination; fully reallocated to audited backlog features.
- Upside ($450k): 0.70 FTE per squad × 4 squads = 2.8 FTE ($560k loaded). 80% elimination with high opportunity realization.
2. Risk Reduction
Centralization standardizes safety filters and evaluation assertions, reducing localized failure rates—but it introduces shared failure risk. The formula nets both.
$$\Delta R_{\text{risk}} = E[L_{\text{local baseline}}] - E[L_{\text{shared platform}}] - E[L_{\text{correlated failure}}]$$
In the base case: each squad has a 20% annual probability of a formatting failure or hallucination reaching customers at $150k remediation cost. The platform reduces that to 3%. But a misconfigured gateway rule now impacts all four squads simultaneously at 4% annual probability and $450k systemic cost.
E[L_local]= 4 × 0.20 × $150k = $120k/yearE[L_shared]= 4 × 0.03 × $150k = $18k/yearE[L_correlated]= 0.04 × $450k = $18k/year- Net Risk Reduction (Base): $84k/year
The correlated failure term is not optional arithmetic. A platform that improves per-squad incident rates while introducing a common-mode failure vector has not reduced risk—it has redistributed it.
3. Compliance Exposure
Compliance exposure isolates statutory regulatory penalties and external audit findings, separate from the operational remediation costs modeled in risk reduction.
$$R_{\text{compliance}} = P(\text{statutory breach}) \times \text{Expected Statutory Fine} \times \text{Mitigation Efficacy}$$
- Conservative ($0): In unregulated workflows, this line item is zero. Do not carry it forward.
- Base ($50k): 5% annual probability of a reportable PII de-identification audit finding under HIPAA/GDPR × $1.5M statutory fine × 67% mitigation efficacy via deterministic Presidio masking.
- Upside ($100k): 10% annual audit probability under heightened inspection regimes × $1.5M fine × 67% efficacy.
4. Distillation Savings (Net)
Inference cost reduction achieved by replacing generalist frontier models with task-distilled local models.
For the reference workload (50 billion annual tokens; 60% eligible for distillation; $2.35 delta cost per million tokens):
Gross Savings = 50,000M tokens × 0.60 × $2.35/M = $70,500/year
Subtracting operational costs (fine-tuning, hosting, evaluation, maintenance, fallback)—which equal ~50% of gross savings in Year 1:
- Conservative ($0): Gated to zero until inference volumes and evaluation harnesses mature.
- Base ($35k): $70.5k gross minus $35.5k operational costs.
- Upside ($70k): Full optimization as token volume expands and fixed hosting costs amortize.
5. Churn Protection
- Conservative & Base ($0): Retaining accounts cannot be defensibly attributed to platform infrastructure in early stages. Multi-quarter customer telemetry isolating extraction accuracy from pricing and product ergonomics is required before this line item earns a place in the base case.
- Upside ($65k): Preservation of one enterprise account ($65k ACV) in a squad directly attributable to zero-defect SLA compliance and audit verification.
The Scenario Model
| Benefit Term | Conservative | Base | Upside | Evidence Quality | Confidence Condition |
|---|---|---|---|---|---|
| Avoided rework | $150k | $320k | $450k | High | Audited contractor spend reduction or signed backlog reallocations |
| Risk reduction | $20k | $84k | $134k | Medium | Net of correlated-failure term; requires prior error incident baseline |
| Compliance exposure | $0 | $50k | $100k | Medium | External statutory fines only—no overlap with operational risk |
| Distillation savings | $0 | $35k | $70k | Low–Medium | Net of fine-tuning, hosting, and eval costs; zero in Year 1 conservative |
| Churn protection | $0 | $0 | $65k | Low | Upside only; requires documented customer contract SLA link |
| Total Annual Gross Benefit | $170k | $489k | $819k | — | — |
Three Years of Phased Cash Flows
A flat Year 1 table obscures the true capital profile. Upfront build costs are front-loaded; distillation, learning, and compounding benefits materialize in later years. Below is the multi-year schedule evaluated at a 10% WACC, with a $140k initial build cost and $54k Year 1 / $40k Year 2–3 operating overhead.
Conservative
| Cash Flow Line | Year 1 | Year 2 | Year 3 |
|---|---|---|---|
| Gross Realized Benefits | $170,000 | $190,000 | $210,000 |
| Platform Cost (Cap + Op) | ($194,000) | ($40,000) | ($40,000) |
| Net Annual Cash Flow | ($24,000) | $150,000 | $170,000 |
| Discounted Cash Flow | ($21,818) | $123,967 | $127,724 |
| 3-Year NPV @ 10% WACC | $229,873 |
Base Case
| Cash Flow Line | Year 1 | Year 2 | Year 3 |
|---|---|---|---|
| Gross Realized Benefits | $489,000 | $540,000 | $600,000 |
| Platform Cost (Cap + Op) | ($194,000) | ($40,000) | ($40,000) |
| Net Annual Cash Flow | $295,000 | $500,000 | $560,000 |
| Discounted Cash Flow | $268,182 | $413,223 | $420,736 |
| 3-Year NPV @ 10% WACC | $1,102,141 |
Upside Case
| Cash Flow Line | Year 1 | Year 2 | Year 3 |
|---|---|---|---|
| Gross Realized Benefits | $819,000 | $920,000 | $1,050,000 |
| Platform Cost (Cap + Op) | ($194,000) | ($40,000) | ($40,000) |
| Net Annual Cash Flow | $625,000 | $880,000 | $1,010,000 |
| Discounted Cash Flow | $568,182 | $727,273 | $758,828 |
| 3-Year NPV @ 10% WACC | $2,054,283 |
Two things the phased schedule reveals that the flat table hid:
- Conservative breakeven occurs in Month 14. Year 1 loses $24k by design—initial build is front-loaded. Year 2 and Year 3 capture ongoing rework avoidance and risk reduction without repeating that cost.
- Capital efficiency accelerates. Operating costs drop from $194k in Year 1 to $40k in subsequent years, expanding operating leverage across all three scenarios.
Resolving the Kill Rule vs. Patient Capital Tension
A fundamental tension runs through every AI platform investment:
- The Patient-Capital Claim: AI moats require years of governed production operation, compounding feedback loops, and curated evaluation data to mature.
- The Traditional Kill Rule: Cut funding if the platform does not achieve target adoption and net positive returns after two release cycles.
If conservative returns are negative in Year 1 (−$24k) because distillation and learning benefits require multiple cycles to compound, a rigid two-cycle kill rule will execute the platform precisely when it needs patient capital.
Platform Stages as Call Options
The resolution is to replace point-estimate capital budgeting with real-options valuation.
flowchart TD
INVEST["**Stage 1 / Stage 2 Investment**\nPurchases Call Option on Shared Leverage"]
WINDOW["90-Day Evidence & Option Window"]
INVEST --> WINDOW
WINDOW --> ALIVE
WINDOW --> KILL
ALIVE["**Uncertainty Unresolved**\n• Schema overlap promising\n• Active squad feedback\n• Low marginal pod cost"]
KILL["**Hypothesis Invalidated**\n• Squads bypass platform\n• Schemas structurally diverge\n• Overlays exceed shared core"]
ALIVE --> EXTEND["**Keep Option Alive**\nExtend Stage 2 Pod\nDo NOT advance to Stage 3"]
KILL --> EXIT["**Exercise Kill**\nDecommission contract\nAbsorb reversal cost early"]
style INVEST fill:#1e293b,stroke:#60a5fa,stroke-width:2px,color:#f8fafc
style WINDOW fill:#1e293b,stroke:#94a3b8,stroke-width:1px,color:#cbd5e1
style ALIVE fill:#064e3b,stroke:#10b981,stroke-width:1.5px,color:#f8fafc
style KILL fill:#450a0a,stroke:#f87171,stroke-width:1.5px,color:#f8fafc
style EXTEND fill:#065f46,stroke:#34d399,stroke-width:1px,color:#f8fafc
style EXIT fill:#7f1d1d,stroke:#fca5a5,stroke-width:1px,color:#f8fafc
Stage 1 and Stage 2 are call options, not full infrastructure commitments. They purchase the right—but not the obligation—to expand to a Stage 3 shared service once uncertainty resolves. The value of the pilot is information value: confirming whether squads share evidence semantics, verifying operator override volume, and testing squad collaboration.
The options-aware kill and extension rule:
- Keep the Option (Maintain Stage 2): If adoption is slower than projected but squads actively use the capability contract, feedback is accumulating, and semantic divergence is low—do not kill. The uncertainty has not resolved. The option to expand remains valuable.
- Exercise the Kill (Contract to Local): If squads actively bypass the platform, domain overlays exceed the shared core (Reuse Trap), or automated telemetry shows zero operator correction engagement—the hypothesis is invalidated. Kill immediately. Revert to local squad ownership before reversal costs compound.
Where This Approach Fails
A decision framework that does not identify its own failure conditions is marketing copy.
1. Workload Scarcity. One or two AI features with low transaction volume cannot amortize platform build and coordination costs. The conservative scenario produces negative net returns. Do not build a platform. Revisit when volume and semantic overlap mature.
2. The Broken Learning Flywheel. Distillation depends on an operational chain: production operation → operator override → validated correction → curated evaluation corpus → distilled model. If operators correct inconsistently, enterprise customers restrict cross-tenant data pooling, or sparse overrides fail to reach statistical significance, the chain breaks. Gate distillation savings to the upside scenario only until automated telemetry verifies closed-loop curation in production.
3. The Moving Model Frontier. Frontier model capabilities improve materially every six to twelve months. A platform that encodes custom prompt patterns, bespoke routing heuristics, or fragile retrieval wrappers finds that an upstream model release renders eighteen months of internal engineering obsolete. Decouple platform contracts from model implementation details. Anchor leverage in proprietary organizational state and task-specific evaluation suites.
4. The False Reuse Trap at Scale. Centralization gates are relaxed under organizational pressure, grouping superficially similar workflows into a shared abstraction. The platform team spends its roadmap maintaining dozens of configuration flags and divergent assertion suites for incompatible squads. Enforce the 2-Week Convergence Spike and continuously track the reuse ratio. Decommission shared contracts when domain overlays exceed common logic.
5. Governance Gridlock. Centralization concentrates institutional risk, attracting multi-stakeholder oversight. Every squad release requires sequential approvals from model risk governance, data privacy, security architecture, and the platform team. Technical fragmentation is replaced with organizational paralysis. Automate governance controls directly at the platform gateway layer, tier controls by risk classification, and preserve autonomous self-service for low-risk workflow iterations.
6. The Compliance Blanket Failure. Treating regulatory compliance as an unchallengeable blank check bypasses financial scrutiny. CFOs audit compliance-driven platform expenditures and demand verifiable proof of control efficacy. Underwrite compliance defensively using actuarial expected loss ($P(\text{Breach}) \times \text{Fine} \times \text{Efficacy}$) rather than qualitative assertions.
How Product Managers Know the Platform Is Working
A platform that cannot be measured cannot be governed.
| Metric Category | Metric | Target Threshold |
|---|---|---|
| Adoption | Workflow adoption rate | ≥ 75% of eligible squads by month 6 |
| Adoption | Platform bypass rate | < 5% |
| Efficiency | Time to first capability invocation | < 3 business days from squad onboarding |
| Efficiency | Contract-change lead time | < 5 business days |
| Quality | Evaluation coverage | 100% of production inference volume |
| Quality | Escaped regression rate | < 1% of deployments |
| Economics | Realized unit cost | Downward trend Q/Q |
| Learning | Correction retention rate | ≥ 40% of overrides promoted to eval suites |
| Learning | Reuse ratio | < 0.25 (above signals Reuse Trap) |
| Governance | Domain-pod autonomy score | ≥ 90% of pod releases without central approval |
| Reversibility | Decoupling lead time | ≤ 10 engineer-days per consuming squad |
Watch for four tripwires that signal the platform has expanded beyond its demonstrated reuse boundary:
- Adoption collapse: Workflow adoption drops below 75% or bypass rises above 5%.
- Reuse Trap breached: The reuse ratio exceeds 0.25.
- Broken learning pipeline: Correction retention falls below 40%.
- Architecture rigidity: Decoupling lead time exceeds 10 engineer-days.
The Decision Gate: Three Outcomes
| Outcome | Conditions | Action |
|---|---|---|
| Proceed | All required centralization gates pass. Three or more scored dimensions demonstrate meaningful overlap. Base-case scenario produces positive 3-year NPV against the best local alternative. | Advance to next platform stage. Assign capability steward. Establish operating metrics and option review milestones. |
| Pilot | Required gates pass but scored overlap is borderline, or the base-case ROI requires validating assumptions through a narrow staged release. | Advance to a time-limited Stage 2 pilot with two squads. Fund the 2-week Convergence Spike. Price as a call option with a 90-day review gate. |
| Remain Local | A required gate fails, scored overlap is structurally divergent, or the conservative scenario does not justify investment without an unvalidated compliance excuse. | Retain capability inside the product squad. Continue managed model access. Re-evaluate when volume, reuse, or risk classification warrants a new assessment. |
One-Line Synthesis
The platform must earn the right to expand—from managed access, to domain capability, to shared service—through demonstrated reuse, measurable outcomes, and realized economics. Sequence capital on evidence, price early stages as real options, and never confuse architectural enthusiasm with financial return.
Related Reading
- Before You Fund an AI Platform: Prove the Workflow First — Part 1: The capability sensitivity test and sequencing thresholds.
- What to Build, Buy, and Federate in an Enterprise AI Platform — Part 2: The architecture ownership decision and the three-tier federated model.
- The AI Pricing Paradox: Why Cost-Plus SaaS Models Collapse — How probabilistic compute costs destabilize traditional subscription pricing.
- EROI Is Not Enough: The Capital Return Trap of Generative AI — Why traditional return on investment metrics miss the systemic depreciation of generative software assets.
- Policy-as-Code for Autonomous Agents — Implementing deterministic verification perimeters around probabilistic reasoning engines.
The ideas in this post are my own — they emerged from questions I asked while learning applied AI concepts and putting them to work in my job and my projects. The prose was developed with AI assistance.
Frequently Asked Questions
What is the correct baseline for an AI platform ROI calculation?
The baseline for platform ROI is not 'squads doing nothing' or 'squads building everything from scratch poorly.' The correct baseline is the best realistic local alternative: product squads buying managed model access, retaining local capability contracts, and independently handling governance. Platform incremental value is the difference in cost, risk, and learning outcomes between the federated platform model and that best local alternative.
How do you resolve the tension between the Kill Rule and patient capital in AI platforms?
Frame staged platform investment as purchasing a real option on enterprise leverage rather than a point-estimate NPV commitment. Early stages purchase information value—reducing uncertainty about cross-squad semantic overlap and adoption friction—before committing to heavy shared service infrastructure. Replace blunt time-based kill gates with an options-aware rule: if uncertainty remains unresolved but squads actively collaborate, extend the low-cost domain pod stage; if the hypothesis is invalidated through contract bypass or structural schema divergence, exercise the kill.
Download the Architecture of Proof Checklist
Ready to implement? Get the definitive checklist for building verifiable AI systems.