CFO guide to AI agents: fund the work you can verify

Time saved is not the same as cash saved. Learn to assess AI agent costs, validate business value, and decide whether to scale a pilot.

Time saved isn’t the same as cash saved. Learn to assess AI agent costs, validate business value, and decide whether to scale a pilot.

A pilot can save time without earning a larger budget. Your funding decision depends on what happened to that time, the work, and the spending.

The demo may show completed conversations. Operations may report shorter handling times. Neither tells you whether the organization accepted the result or avoided an actual expense.

A useful CFO guide to AI agents starts with a narrower question: what evidence would justify the next commitment? You need a defined workflow, a credible comparison, full costs, and someone accountable for the result.

That evidence can support expansion, a smaller experiment, a narrower task, or a stop. The point is to distinguish those decisions before the annual forecast makes them look inevitable.

TLDR

  • Define the business result precisely and include the work people must finish when the agent cannot.
  • Separate delivery costs, cash changes, and noncash value; keep past spending distinct from the next commitment.
  • Fund the rationale you can support, whether that is savings, service quality, growth capacity, or reduced risk.

Start with the decision Finance must make

Your AI investment business case should name the commitment before presenting the forecast. Are you approving another test, more eligible work, or a production service with ongoing obligations?

Those decisions need different evidence. A short extension might resolve uncertainty about reviewer effort. A larger deployment requires confidence in operating costs and accountability for failures.

Take a service-response pilot: the agent retrieves records and prepares drafts for human review. You’re deciding whether to expand that workflow, not whether every customer case can now resolve itself.

A sourcing decision already exists. Reopening the entire build-versus-buy decision would distract from the question facing Finance now.

Instead, write a conditional decision: expand if the proposed workload meets the agreed quality, service, cost, and control requirements. Then identify which investment rationale must hold.

The finance partner may accept the unit cost but reject the cash assumption. The service owner may accept the economics but question the workload exclusions.

A useful memo shows both objections. A blended confidence score can hide them.

Define the work you will accept and compare

Choose one outcome and use it throughout the baseline, costs, and benefits. An accepted draft, a sent response, and a resolved customer case are different units.

This article’s example uses an accepted response draft: one draft for a unique request that passes human review within the agreed service window.

Acceptance requires the correct account, current supporting evidence, a supported answer, and appropriate wording. The reviewer can accept it for the next communication step without further drafting work.

That doesn’t mean the response was sent or the customer’s underlying problem was resolved. Those later stages remain outside this example and must be treated consistently in both alternatives.

Count the business result once

Separate eligible requests, agent attempts, accepted drafts, and fallback work. Several model calls and retries can contribute to one accepted draft; they don’t create several outcomes.

The FinOps Foundation’s token-economics guidance states: “Failed and discarded runs are cost, not completions.” Its outcome measure also requires a workload-specific quality floor.

Classify accepted drafts by how they were completed. Agent-assisted work and human fallback can satisfy the same business requirement, but they aren’t evidence of the same automation rate.

Define how you handle corrections after acceptance. Keep the review window consistent so later rework does not disappear between reporting periods.

Assign the work the agent does not finish

A rejected draft still leaves a customer request in the queue. Name who fixes it, who handles escalations, and where that work’s cost belongs.

In this example, a service team rewrites or completes drafts the agent cannot deliver. Routine review and additional fallback drafting have separate cost categories.

Unresolved requests stay on the workload report with an owner and expected completion window. Costs incurred so far remain visible; future completion costs need a separate forecast, not an invented success count.

If the comparison assumes all requests finish within a service window, unfinished requests mean that assumption failed. Do not declare a cheaper equivalent service by dropping the backlog.

Match quality and service levels

Your counterfactual describes what would happen without the proposed investment. Compare the same request mix, volume, acceptance requirements, service window, and downstream work.

A reviewed agent-assisted draft should be compared with a reviewed human draft, not an unreviewed answer or a fully resolved customer case.

Use a matched comparison or controlled rollout where appropriate. Disclose differences in complexity, seasonality, or exclusions rather than hiding them in an average.

Operations owns the definition of acceptable work. Finance agrees how benefits will be recognized. Platform owners identify resources consumed across successful and unsuccessful attempts.

What is cost per accepted outcome?

Cost per accepted outcome divides a workflow’s in-scope cost by its unique qualifying results over the same period. Here, the result is an accepted response draft. Include failed attempts, required review, and human fallback in the cost. The measure describes delivery economics, not whether cash spending fell.

Recurring cost per accepted draft = recurring workflow cost ÷ unique accepted drafts

The FinOps unit-economics capability connects expenditure with meaningful units while balancing cost, quality, speed, and risk. The useful unit depends on the work you are funding.

If another report says “cost per successful outcome,” check its acceptance rule before comparing figures. The label alone does not establish equivalent work.

Include runtime, infrastructure, licenses, verification, fallback, and ongoing engineering or security work where applicable. Give every cost one home.

If runtime includes failed calls, don’t add the same charges again as a failure allowance. Human rework needs its own measured or disclosed cost basis.

For detailed inputs, use the AI agent total cost of ownership breakdown. The AI agent cost management guide covers the measurement method and operating cadence.

If no drafts qualify, unit cost is undefined. Report the spending and unfinished work rather than presenting zero as an attractive result.

Which benefits justify the investment?

Cash savings are one valid rationale, not the only one. Service quality, growth capacity, and risk reduction can also support funding when their evidence and costs are explicit.

Cash savings require a changed outlay

Released capacity changes cash spending only when it changes an actual outlay against a credible comparison. An employee using saved time to address a backlog hasn’t made their salary disappear.

HM Treasury’s Government Efficiency Framework, updated in 2025, distinguishes reduced expenditure from productivity gains without reduced spending. Its scope is UK central government, not corporate accounting law.

Ask your finance team to establish the company’s realization rule. A smaller cancellable supplier bill and an unchanged salary allocation should not enter the cash model as equivalent benefits.

Noncash value still needs an owner and evidence

Better service might mean more drafts accepted without correction or more requests handled within the agreed response window. Measure those outcomes against the current process rather than attaching an arbitrary dollar value.

Growth capacity can justify investment when the team can absorb an evidenced workload increase. Show demand assumptions, remaining bottlenecks, and the work people will do with the capacity.

That’s different from an avoided hire. The latter requires a credible hiring counterfactual and a supported change in future spending.

Risk reduction can also matter, but it needs evidence of the specific failure becoming less likely or less consequential. Do not manufacture an expected-loss figure to make the spreadsheet positive.

The memo can recommend investment for one of these reasons while showing negative near-term cash. State the trade-off and the threshold the accountable owner will use.

Work through a funding example

Every input below is hypothetical, not observed pilot data or a customer result. The example compares two ways to deliver accepted response drafts, not sent responses or resolved cases.

Assume 1,000 eligible requests arrive each month. The agent-assisted path produces 800 accepted drafts; the service team completes the remaining 200 through fallback drafting and review.

All 1,000 meet the same acceptance requirements and service window. The 200 fallback requests remain visible in the operating report, even though their drafts ultimately qualify.

The baseline supplier would deliver 1,000 equivalent accepted drafts for $8,000 per month. Assume its charge includes the same drafting and review requirements, and can actually be canceled.

Hypothetical monthly inputAssumed amountScope
Runtime, infrastructure, and licenses$2,400All successful, failed, and retried attempts, counted once
Routine verification$2,000Review across all 1,000 drafts; excludes additional fallback drafting
Additional failure handling and fallback$1,200The service team’s work to complete the remaining 200 drafts; excludes routine review
Ongoing integration, security, and operations$800Recurring support, separate from past pilot setup
Total recurring workflow cost$6,400Four non-overlapping categories
Accepted response drafts1,000800 agent-assisted plus 200 human fallback
Recurring cost per accepted draft$6.40$6,400 ÷ 1,000
Comparable baseline cost per accepted draft$8.00$8,000 ÷ 1,000
Assumed monthly spending avoided$8,000Full supplier bill for the equivalent workload, once cancellation takes effect

In short: the comparison includes human completion of the difficult work rather than pricing only the agent’s successful subset.

For this cash illustration, assume every recurring category is an incremental cash outlay. If your figures contain unchanged employee salaries, separate the economic-cost and cash views first.

Replace the fallback estimate with measured handling time and an appropriate hourly cost, including escalation work. Keep it separate from routine review.

Sending responses and resolving customer cases remain outside both alternatives. If either path changes those downstream costs or service outcomes, extend the comparison before approving it.

Separate money already spent from the next commitment

Assume the pilot’s $9,600 setup and testing payment has already been spent. This simplified example assumes no other historical pilot cash flows and no new expansion setup payment.

Keep that $9,600 in the full-program view. Do not charge it again as cash required for the next decision.

Hypothetical viewCalculationResult
Monthly net cash after supplier cancellation$8,000 − $6,400$1,600
Next 12 months’ incremental cash outlays12 × $6,400; no new setup assumed$76,800
Next 12 months’ net cash versus baseline12 × $1,600$19,200
Full-program net cash through that period$19,200 − $9,600 already spent$9,600
Cost including past setup and the next 12 months($9,600 + 12 × $6,400) ÷ (12 × 1,000)$7.20 per accepted draft

In short: the program’s historical performance and the cash needed from today are different views of the same decision.

At $1,600 monthly net benefit, six months would recover the $9,600 historical spend after benefits begin. That is historical cost recovery, not a fresh setup payment or a reason to scale by itself.

If expansion requires new configuration, migration, training, or contractual payments, add those future amounts to the next-commitment view. A sunk cost does not make a weak future investment attractive.

These calculations assume steady volume and costs, immediate supplier cancellation, and no benefit delay. They omit discounting, tax, working-capital effects, and terminal value; they are not NPV or IRR.

For broader AI agent ROI mechanics, use the AI ROI guide. This memo needs a decision about the actual commitment, not another return label.

What would change the decision?

Test volume, completion effort, benefit timing, and commitments before approving expansion. Start with changes you can interpret separately, then consider combined downside.

Unit cost rises to $8 if only 800 drafts qualify at the same $6,400 cost. The missing 200 also create a service and completion-cost problem.

Identify who finishes them, forecast the additional cost, and re-estimate the comparable benefit. Do not hold the full supplier saving constant if the replacement no longer delivers equivalent work.

If realizable monthly cash benefit falls to $6,000 at unchanged recurring cost, monthly net cash becomes negative $400. The cash-saving rationale fails under that assumption, even if a separate service rationale remains worth reviewing.

Model delayed benefits and committed costs

Suppose supplier cancellation takes effect only in month three, while the new workflow incurs its full monthly cost from month one. The next-year net cash becomes:

10 × $8,000 − 12 × $6,400 = $3,200

Including the $9,600 already spent leaves full-program net cash at negative $6,400 through that horizon. The timing change matters even though the eventual monthly economics are unchanged.

Check minimum terms, committed usage, termination fees, and overlapping service periods. A stop decision may remove variable spending without eliminating contractual payments already committed.

Show what is avoidable from the decision date and what remains payable. Do not treat a nominal monthly price as a cancel-anytime arrangement without checking the contract.

Put the evidence into a funding memo

A practical CFO guide to AI agents ends here: use one page to connect the requested commitment with its evidence and conditions. Missing information is a valid entry when it has an owner and changes the decision appropriately.

Memo fieldWhat to record
Decision requestedScale, bounded extension, narrower scope, or stop; specify new spending and commitments
Business unitOne accepted result, its quality criteria, service window, and downstream exclusions
Complete workloadAgent-assisted results, human fallback, unresolved work, and completion owners
Comparable baselineEquivalent volume, complexity, quality, service level, and remaining human work
Evidence statusObserved, assumed, or missing for each input, with a source and date
Cost viewsHistorical spend, future incremental outlays, unavoidable commitments, and full-program economics
Investment rationaleCash, service quality, growth capacity, or risk reduction; do not blend their claims
DownsideLower acceptance, more fallback, benefit delays, workload changes, and exit costs
Decision conditionThe evidence or threshold needed for the next commitment
OwnershipFinance and operating owners, review date, and action if conditions fail

In short: the memo makes the next commitment conditional on evidence, rather than on defending the pilot’s past spending.

Scale when the evidence supports the expanded workload and its operating conditions. Do not assume the next request mix resembles the selected pilot.

Extend when a bounded experiment can answer a specific uncertainty. Give it a budget, a question, and an end date.

Narrow when value depends on a better-defined scope. A workflow may handle current documentation well but struggle with disputed records.

Stop when the future commitment is no longer justified or an unacceptable control failure requires intervention. Include remaining contractual and service obligations in the exit decision.

For the wider committee’s responsibilities, use the enterprise AI agent buying criteria. Keep this memo focused on what Finance must approve now.

Apply the same definitions of work, timing, and evidence when answering each question.

If you’re evaluating Computer, by DevRev, bring this memo to the pilot discussion. Define the work you expect it to support, then test acceptance, human fallback, total cost, and the benefit you intend to realize.

Approve the next commitment, not the most optimistic spreadsheet. A product demonstration can start the conversation; evidence about your workflow should decide the funding.

Frequently Asked Questions

DEVREV

See Computer work for you

Your AI teammate that finds answers, takes action, and gets work done across every tool.