AI Process Integration ROI: the analysis

What the evidence says, how companies decide what to automate, and how to measure the result · analysis as of September 18, 2026 · case data refreshes weekly on the library page

Executive summary Evidence bar State of adoption How companies decide What the cases show Measuring ROI Keeping it current

Executive summary

Measurable AI ROI is real but concentrated: roughly 88% of large organizations now use AI somewhere, yet only 37% can attribute any EBIT impact to it and only about 6% attribute 5% or more (McKinsey, Aug 2026). BCG's 2025 survey of 1,250 executives puts 60% of companies in a "laggard" group with minimal revenue or cost gains (BCG), and MIT NANDA's widely cited 2025 review concluded that 95% of organizations were seeing no measurable return from roughly $30–40B of generative AI spend (MIT NANDA).

The companies that do capture value share a recognizable pattern. They pick a narrow, high-volume, rules-plus-judgment process with a clean baseline (invoices per day, days sales outstanding, tickets per agent, hours per close), redesign the workflow rather than layering a chatbot on top of it (73% of McKinsey's high performers redesigned workflows vs. 25% of everyone else), and measure against the pre-AI baseline for at least a quarter. Back-office finance and operations processes produce the most reliable quantified returns; MIT NANDA found budgets skewed toward sales and marketing "despite better ROI in operations and finance."

The strongest named-company evidence clusters in five process families: invoice and accounts-payable automation (60–90% touchless rates, 65–87% time reduction per document), receivables and collections (5.5–28% DSO reductions, $20M+ cash-flow gains at industrial distributors and CPG firms), expense audit (63–74% auto-approval, 2,000–16,000 audit hours saved per year), customer-support agent assist (14% more issues resolved per hour in the only peer-reviewed field study; 30–49% reductions in after-call work or transfers), and software maintenance (Amazon's $260M annual savings from automated Java upgrades). Headline claims about headcount, by contrast, have a poor track record: Klarna, Commonwealth Bank, and Duolingo each walked back or reframed AI-for-people announcements within a year.

The practical reading for a $50M+ company: treat AI as process re-engineering with a measurement plan, start with a back-office process where the unit cost is already known, buy rather than build for the first wave (MIT NANDA found vendor tools succeed about twice as often as internal builds), and expect a 3–12 month payback on the good cases and nothing at all on pilots without a baseline.

Evidence bar and how to read a case

A case enters the library only if a named company (or a rigorous study) states a quantified before/after outcome for a specific business process. Anonymous "a Fortune 500 firm" claims and vendor aggregates without a named customer are excluded, with the exception of peer-reviewed studies that anonymize the firm by design.

TierMeaningTypical source
HighPeer-reviewed or randomized field study, or a company figure reproduced consistently over timeNBER, QJE, Management Science, HBR case co-authored with the company
Medium-highNamed company, named executive, stated baseline and timeframe; vendor may host the storyMicrosoft/AWS customer stories with baselines, company press releases with methodology
MediumNamed company and metric but missing baseline or timeframe; self-reportedVendor case studies, earnings-call remarks, CEO interviews
Low–mediumPilot-stage, projected, or survey-of-users figures"Forecast benefit", "up to" claims, POCs

Three reading rules apply. "Vendor-reported" means the number comes from the software vendor's marketing page, even when a customer executive is quoted; these figures are directionally useful but rarely audited. Capacity claims ("equivalent to 700 agents", "4,500 developer-years") are estimates against a counterfactual and are not the same as booked savings. A company's later reversal or clarification is recorded with the case rather than deleting it, because the reversal is itself evidence about what the process change did and did not deliver.

State of adoption and returns

Adoption is near-universal among large companies while measured financial impact has stalled at roughly a third of them for two years running. The large surveys agree on the shape even though they disagree on the numbers, because each samples a different population: McKinsey and BCG survey executives at mostly $1B+ firms, the U.S. Census surveys all employer businesses, and MIT NANDA reviewed disclosed initiatives plus interviews. The full macro table, with sample sizes and links, is on the library page.

Three conclusions follow. Individual-productivity deployments (a Copilot or ChatGPT seat per employee) reliably produce self-reported time savings of 20–45 minutes a day but rarely show up in EBIT, which is the gap between McKinsey's 80% "AI improved my productivity" and 37% "EBIT impact". Process-level deployments with a redesigned workflow are what separate the 5–6% high performers from the rest. And the failure mode is not the model but the absence of a baseline, an owner, and a mechanism for the system to retain feedback and improve; MIT NANDA identified "learning rather than infrastructure, regulation, or talent" as the core barrier.

A note on the 95% figure: it comes from a small, non-random sample (52 interviews, 153 conference surveys) and defines "return" as measurable P&L impact within about six months of the pilot. It is best read as a statement about pilot design rather than about the technology.

How companies decide what to automate

The companies with quantified results chose processes by unit economics first and technology second. The common screen, reconstructed from the cases and from the BCG, McKinsey and MIT NANDA analyses, runs on five tests.

  1. Volume and repetition: the process runs thousands of times a month with a known cost per unit (GameStop hand-keyed 750,000 invoices a year; Danone handled 1.1M deduction claims; Klarna had 2.3M support chats a month).
  2. A measurable baseline already exists: cost per invoice, DSO, first-contact resolution, hours per close, time-to-hire. If the metric has to be invented for the pilot, the pilot will not prove ROI.
  3. Structured judgment, not open-ended creativity: matching, classifying, extracting, summarizing, routing and drafting against rules. The academic evidence shows the largest gains where the task sits inside the model's competence and the smallest, or negative, gains outside it.
  4. Tolerable error cost with a human backstop: exception queues, auto-approval thresholds, or agent-assist rather than agent-replace.
  5. A named owner in the function, not in IT: BCG's 70% "people, process and behavior" share is the reason, and Deloitte finds only 30% of firms actually redesign the process.

Screening flow

1
Inventory processes by volume × cost per unit.
2
Is the baseline measurable? No → instrument first, revisit next quarter.
3
Is it a bounded rules-plus-judgment task? No → human-led, AI assist only.
4
Is the error cost tolerable with a backstop? No → human-led, AI assist only.
5
Buy or embed in the existing system; run a 90-day pilot against the baseline with an owner in the function.
6
Payback under 12 months? Yes → scale and redesign the workflow. No → back to step 2.

Build, buy or embed. The evidence favors buying or embedding for a first wave. MIT NANDA found vendor-built tools succeed roughly twice as often as internal builds; the finance cases with the cleanest numbers all run on embedded vendor AI (Coupa, HighRadius, AppZen, Rossum, Vic.ai, Pactum) inside an existing ERP or P2P flow. Internal builds pay off where the company has a large engineering organization and a proprietary data advantage (JPMorgan's LLM Suite, Airbnb's test-migration pipeline, Amazon Q for Java upgrades). McKinsey's 2026 survey notes nearly a third of respondents declined a software purchase in favor of building with agentic coding tools, so the line is moving, but the ROI evidence for internal builds is still concentrated in software-heavy firms.

Where budgets go versus where returns are. MIT NANDA reports AI budgets skewed to sales and marketing while back-office operations and finance show better ROI; BCG puts 70% of AI value in core functions such as sales, manufacturing, supply chain and pricing and only 13% in IT. Read together: front-office revenue lift is larger but harder to attribute, back-office cost reduction is smaller but provable. A company that needs to demonstrate ROI to a board should start in the back office and fund front-office experiments from those savings.

Governance. Deloitte finds only 21% of organizations have a mature governance model for agents, and Gartner expects 40%+ of agentic projects to be canceled by end-2027 on cost, unclear value or inadequate risk controls. The practical minimum is a use-case register with owner, baseline metric, target, decision date and kill criteria, reviewed quarterly.

What the cases show, by function

Finance and back office. Touchless rates of 60–90% and per-document time reductions of 65–88% are repeatable across AP vendors and industries once invoice volume exceeds roughly 100,000 a year. Receivables AI pays back in cash rather than headcount: DSO reductions of 5–28% and recovered deductions in the tens of millions at CPG scale. Expense audit is the cleanest "100% coverage replaces sampling" story, with 63–74% auto-approval and 2,000–16,000 hours a year saved. Close and reconciliation gains are real but small and mostly hours-based. Tax and external audit have no named-company quantified cases yet.

Customer-facing. The rigorous number is +14% resolved issues per agent-hour, concentrated in less-experienced agents, which shows up as slower hiring rather than layoffs. Company-reported containment of two-thirds of chats (Klarna) or 70% first-time resolution (Vodafone) is achievable for high-volume, low-complexity contact types, but the Klarna and CBA episodes show that removing humans faster than the containment rate justifies produces quality and volume rebounds within 12 months. Revenue-side claims (Amazon's $10B, Lumen's $50M) rest on attribution models and belong in a lower confidence tier than cost-side numbers.

Operations and energy. Quantified evidence exists for negotiation (Walmart) and fulfillment automation (Amazon), not yet for AI route or load optimization at a named shipper. The DeepMind cooling result remains the only named, quantified AI energy-cost case at scale; 10–15% of controllable HVAC energy is the defensible planning figure, and a vendor should name a reference site with metered before/after data.

People and IT. Bounded service-desk automation (IBM, Bank of America) delivers 50–75% ticket or call reductions because the question set is finite. Recruiting automation compresses cycle time (Chipotle's 12 → 4 days) because the bottleneck was scheduling. Software maintenance at scale (Amazon, Airbnb) is the clearest "millions of dollars, one project" story. Seat-based copilots deliver 20–45 self-reported minutes a day, which only becomes ROI if the organization redesigns work to capture it; the Danish study's 2.8% time saving and zero wage effect is the honest base rate for copilots without process change.

Measuring ROI

The cases that survive scrutiny all measure one process-level metric against a pre-AI baseline over a fixed window, then convert it to dollars with a rate the finance team already accepts. The cases that get reversed measure capacity ("equivalent to N people") and act on it before the quality metrics catch up.

ProcessPrimary metricGuardrail metricDollar conversion
Accounts payableCost per invoice; touchless rateException rate; duplicate/late payments(baseline − new cost) × annual volume; discounts captured
Accounts receivableDSO; cash-application auto-match rateDisputed-invoice rate; write-offsDSO days reduced × daily revenue × cost of capital
Deductions and claimsDays to resolve; % invalid deductions recoveredDisputes reopenedRecovered dollars net of tool cost
ProcurementSavings % on negotiated spend; cycle time to POSupplier acceptance; maverick spendSavings % × addressable spend
Expense auditAuto-approval rate; audit hoursOut-of-policy catch rateHours × loaded rate + recovered spend
Close and reconciliationDays to close; hours per reconciliationPost-close adjustmentsHours × loaded rate
Customer supportIssues resolved per agent-hour; containmentCSAT; repeat contacts; escalationsAvoided hires or BPO spend
RecruitingTime to hire; application completion90-day retentionVacancy days × daily cost of vacancy
IT / HR helpdeskTickets or calls deflectedReopen rate; satisfactionDeflected tickets × cost per ticket
Software maintenanceEngineer-hours per migrationDefect rate post-changeHours × loaded rate
Energy / facilitieskWh per unit of output or square footComfort and fault incidentskWh saved × tariff

Baseline method. Record at least three months of the primary and guardrail metrics before the change, freeze the definitions, and run the AI process on a subset (a business unit, a supplier tier, a contact type) alongside the unchanged remainder for a quarter. The academic studies that produced credible numbers all used a comparison-group design; a company can do a simpler version by phasing rollout across sites or teams.

Common measurement mistakes: counting capacity as savings before headcount or BPO spend actually changes; ignoring guardrail metrics (Klarna and CBA acted on containment forecasts before CSAT and volume data confirmed them); reporting only the winners (the METR study found expert developers 19% slower while believing they were faster); omitting the cost side (McKinsey finds one in five organizations constrained by AI operating costs); and letting the pilot define the metric.

Keeping it current

The case library on the main page refreshes automatically each week: a scheduled research pass checks the recurring sources (McKinsey, Stanford AI Index, BCG, Deloitte, Gartner, U.S. Census BTOS, Ardent Partners, APQC, NBER/SSRN, cloud-provider and vendor case hubs, and finance/CX trade press), adds cases that meet the evidence bar, re-checks High and Medium-high cases for reversals, and records every change in the library's change log. This analysis page is revised by hand when the findings shift; its date is in the header. Current evidence gaps: tax, external audit, AI route/load optimization at a named shipper, and utility-cost AI beyond the DeepMind case.