Business

Measuring AI ROI: What IT Teams Can Actually Show the CFO

Tobias Walsh
Back to Blog Measuring AI ROI for IT teams

At some point, every IT leader who has been approving AI tool budgets gets a question from the CFO that sounds simple: what are we getting for this spend? The question deserves a real answer, and right now, most IT teams do not have one ready.

The gap is not that AI tools are not producing value. Most of them are. The gap is that the value being produced is mostly attributed to productivity changes that are hard to measure cleanly, while the cost is visible and concrete. Finance sees the invoice. Finance does not see the engineering hours saved. That asymmetry makes IT look like it is spending without accountability.

This post is about building a monthly AI spend report that finance can actually parse, and about being honest about which parts of the ROI case are measurable and which require a different kind of argument.

Start with what is already measurable

Before you can show ROI, you need spend data that is complete and attributed. Most IT departments cannot produce this today without significant manual reconstruction. They have invoices from multiple LLM vendors, they have team credit card statements, and they have some developer accounts that nobody in IT approved. Getting to a clean monthly spend number requires consolidating all of those sources.

Once you have consolidated spend data, the first useful thing you can show finance is not ROI at all. It is visibility. A report that breaks down AI spend by team, by month, and by model family is already more information than most companies have. It shows that IT has a handle on the cost. That in itself changes the conversation. Finance is not primarily concerned that AI costs money. Finance is concerned that AI costs money and nobody in IT knows how much or who is spending it.

The per-team attribution view is also where you start finding anomalies. A team that represents 8 percent of headcount should not be running 40 percent of AI spend unless there is a very specific reason. Per-attribution data makes those outliers visible and gives you something to investigate rather than defend.

The cost-side metrics that are actually credible

Several cost savings from AI deployments are measurable with reasonable confidence. They are not always large, but they are defensible.

The first is vendor consolidation. Before a routing layer, teams often maintain direct relationships with multiple LLM vendors because switching costs are high. When routing is centralized, you can see exactly how much of your traffic could route to lower-cost models without changing output quality your teams notice. The difference between routing 70 percent of requests to GPT-4 by default versus routing based on task complexity is real and quantifiable. If you have request logs with costs, you can build a counterfactual: what would last month's traffic have cost under the current routing policy versus the old one?

The second is contract leverage. Dispersed spend across multiple teams has no negotiating weight. Consolidated spend data lets you approach a vendor with an actual volume number. A company spending $120,000 per year spread across 15 team accounts is not a meaningful customer to OpenAI. A company that routes through a central endpoint and can report $120,000 in annual spend with predictable growth is a conversation with a named account rep. The discount range for volume commitments in this market varies but is typically 10 to 30 percent at meaningful scale.

The third is the elimination of spend on decommissioned or redundant agents. When you run a quarterly review of agent activity through the routing dashboard, you reliably find agents that were built for a project that ended and are still processing requests. Turning those off has an immediate and provable cost impact.

Why productivity claims are hard to make cleanly

Here is the honest version of the ROI conversation: productivity gains are real and often substantial, but they are very difficult to attribute to a specific tool with the rigor that CFOs apply to capital expenditure decisions.

The standard productivity argument runs: an agent handles customer support ticket triage, tickets are routed faster, customers wait less, and the support team handles more tickets per day with the same headcount. All of those things may be true. The problem is isolating which part of that improvement is the agent versus the fact that you also hired a more experienced team lead last quarter and changed the ticket routing process in the same period.

IT teams that try to make large productivity ROI claims without controlling for confounding variables tend to have those claims challenged or dismissed by finance. The result is that the AI program looks like it has weak ROI support when in fact the problem is measurement methodology, not actual value.

A more defensible approach: use productivity arguments qualitatively to justify continuation, and use cost metrics quantitatively to justify the current budget level. If you can show that costs are attributed, under control, and trending in a reasonable direction, you have answered the CFO's actual concern. The productivity story is supporting evidence, not the primary data.

What a monthly AI spend report should contain

The report format that gets traction with finance tends to include five things.

First, total spend for the month with a month-over-month comparison. Is spend growing, flat, or shrinking? If it is growing, is the growth proportional to headcount or revenue growth? If it is growing faster than both, that needs an explanation.

Second, per-team attribution. Which teams are spending what, and how does that compare to last month? Anomalies should be flagged with a note, not left unexplained.

Third, model mix. What percentage of requests are going to high-cost models versus lower-cost alternatives? A shift toward more efficient model routing shows up here as cost improvement without sacrificing capability.

Fourth, budget variance. If teams have been given monthly AI budgets, which teams are over and under? Budget overruns that have an explanation are manageable. Budget overruns that are unexplained signal that cost attribution is not working.

Fifth, a brief narrative section. Two to three sentences on anything material that changed this month and one to two sentences on what IT plans to do about any anomalies. Finance does not need a detailed technical explanation. They need to know that IT is watching the numbers and has a plan when something is off.

The governance posture this enables

Building a clean monthly report is not just a finance deliverable. The discipline of producing that report forces IT to have the right data infrastructure in place. If you can produce the report, it means every agent is routed through a central endpoint with team attribution. It means you have spend visibility at the request level. It means you can answer the CFO's question without three weeks of manual reconstruction work.

That infrastructure is also what enables the cost optimization conversations: model routing decisions, vendor contract consolidation, and identification of redundant agents. The report is the visible output of a governance structure that produces value independently of the report itself.

One thing to be clear about: the goal is not to produce a document that justifies AI spending by making the numbers look favorable. The goal is to produce an accurate, reliable view of what AI costs and where the money is going. Finance's trust in IT's AI reporting is built on consistency and accuracy, not on optimistic framing. A monthly report that occasionally shows a cost problem and explains it clearly is more credible than a monthly report that always shows things trending in the right direction.