The $3M in savings slide is almost always wrong — not because the savings don’t exist, but because the methodology didn’t hold up to scrutiny when anyone actually looked.
I’ve sat in enough enterprise finance reviews — 19 years of them at Lockheed Martin, many directly tied to AI and automation projects I led — to recognize the pattern. A team builds something that genuinely improves a process, they attach a number to the improvement, and then that number gets challenged in a budget review and can’t be defended. The project survives or doesn’t based on whether the finance organization decides to extend credibility rather than because the ROI case was actually sound.
Two Failure Modes Before You Calculate Anything
Underselling is the failure mode that gets less attention. ROI exists — real hours are recovered, real error rates drop — but the team has no rigorous way to measure it, so the project gets framed as qualitative improvement. Qualitative improvements don’t survive budget compression. They get cut when finance needs to find savings and the project has nothing auditable to defend itself with.
Overselling is more visible and more damaging to the broader AI program. A team attributes all observed efficiency gains to the AI system when the gains were actually produced by a combination of the tool, a concurrent process redesign, an analyst who changed how she organized her work, and a quarter with lower transaction volume than the prior year. If you claim credit for all of it and a skeptical CFO asks how you isolated the AI contribution, you don’t have an answer. That moment — one moment — kills credibility for the next five AI proposals your organization brings forward.
The Costs That Disappear From ROI Slides
Every ROI methodology I’ve seen that didn’t survive scrutiny had the same structural problem: it counted only the benefits of the AI and the initial build cost, then stopped. The cost categories that routinely disappear from the slide are annotation and labeling cost for supervised systems, the retraining cadence — a model that needs quarterly retraining has a recurring engineering cost that should appear in the denominator, human review of AI outputs in any high-stakes workflow, and the cost of errors the system makes rather than only the errors it prevents.
That last one is particularly common in automation ROI. A system that processes 10,000 transactions per month with a 0.5% error rate introduces 50 errors per month that someone has to catch and correct. If those errors are in finance or compliance workflows, the correction cost — loaded labor rate × correction time, plus any downstream rework — can meaningfully offset the efficiency gains. An honest ROI methodology puts that number in the model. A credibility-optimized ROI methodology hopes no one asks about it.
How a Defense Finance Automation ROI Number Was Built
The $360K annual cost avoidance figure from a major defense aircraft program withholds automation project held up to finance leadership scrutiny because it was constructed to be auditable, not to be impressive. The methodology: loaded labor rate for the analyst population multiplied by 2,132 recovered hours per year — hours that were timed, not estimated, against a documented pre-automation baseline — minus annual system maintenance cost, minus periodic retraining and validation cost for the Alteryx and Python workflows involved, minus an allocation for edge cases that still required manual intervention.
The methodology was presented visibly, not buried in appendices. Finance leadership could see exactly which assumptions the number depended on and could stress-test any of them. When the program office asked what the number looked like if loaded labor rate was adjusted down by 15%, we had that sensitivity analysis ready. The figure that survived scrutiny was lower than the first draft — and more credible precisely because of that.
Framing ROI by Audience
A single ROI figure presented identically to every audience is a communication failure before it’s a methodology failure. CFOs want payback period and net present value — they’re evaluating capital allocation and need to compare this project against others competing for the same budget. Operational leaders want hours recovered per analyst per month because they’re managing capacity and need to understand what changes for their team on Monday. CTOs want the technical debt comparison — build versus buy, and what the maintenance cost trajectory looks like in years two and three when the team that built the system has moved on.
Each of those framings uses the same underlying data. The work is disaggregating the ROI model into components that answer the specific question each audience is actually asking. Presenting a single number to all three audiences and hoping it resonates is why most AI ROI conversations stall in the room rather than producing decisions.
What Transfers
Financial rigor from defense FP&A is not defense-specific. The discipline of building auditable cost and benefit models — with documented baselines, explicit attribution logic, visible cost categories, and audience-specific framings — applies to any AI project that needs to survive budget scrutiny. The aerospace context made the stakes higher and the review process more formal, but the underlying methodology is the same one any enterprise AI team should be using. The first AI project that can’t defend its ROI doesn’t just lose its own budget. It makes the case against the next five.
→ Building something at the intersection of AI, edge computing, and behavioral science? Let’s connect.
Leave a Reply