A friend of mine reached out asking for help assessing whether he should invest in a “measuring ROI on AI” company. This is the second time I’ve been hit with the “should I invest in an ROI of AI company” question recently, and I’ve discussed this issue with at least a dozen executives at companies over the last few months. “What metrics do you report to the board / executive team to prove AI ROI?” It’s a hot topic. Token-maxxing has ended. ROI-maxxing has begun. Long live the maxxing.
So what is the answer? … beautiful dashboards.
Actually, no. But over the course of a year I’ve settled on an effective blueprint, stemming from a core framework in our book Architected Intelligence.
The first problem is that “AI ROI” combines completely different use cases into one metric. Before measuring anything, you have to disambiguate what the AI is doing. The first step is to separate product, process, and productivity.
Although they do overlap, many have found it to be a consistently useful framework. Yes, productivity can turn into a process, and a process can power a product. But each should be tackled in distinct ways!
Note: I’ve intentionally left out the fourth, platform, because a) it doesn’t apply to most companies, and b) I have a separate upcoming post just for you!
ROI of Product AI Is Mostly Just Product
Let’s start with the easiest bucket.
If AI is embedded in a customer-facing product, it is simply another product cost. It might be significant, volatile, or strange, but it sits right alongside the databases, cloud infrastructure, engineers, support, and anything else required to create, maintain, and serve the product.
Ask what experience you are delivering, whether customers value it, and how it affects revenue, retention, and cost. AI may completely transform the product you offer, but you’re still left asking, “Do the product economics work?”
Companies already do this on a regular basis, so no need to fret. What is the ROI of cutting down a customer’s wait time on each click through edge compute instead of having to route all the computation to one centralized location? Is the extra cost worth it? These decisions are routine.
That covers 90% of the “ROI of product” questions. But LLMs throw in a few extra curveballs you should be aware of:
What’s the risk and impact of the LLM behaving poorly? If the LLM for your taco app expresses the wrong opinion on the Peloponnesian War, what are the repercussions?
How much extra ongoing maintenance cost will the LLM require? Maintaining a current knowledge base requires ongoing effort. Otherwise, you may disappoint your taco customers by promising them that sweet sweet beefy Frito burrito that is no longer on the menu. 🪦
Are your customers humans or bots? This is an awkward one that companies are navigating as we speak, but ultimately providing value to your customer is the goal even if you do have to defend against their bot hitting customer service chat over and over asking for free tacos.
Still, in the end these change the maintenance cost and risk profile, but it does not require reinventing some “entirely new theory of AI value in products!”
Process AI Is Just an Operations Problem (this time, with more steps!)
The process bucket also shouldn’t be too scary.
Imagine a defined input, a standard operating procedure, and an expected output.
Complaint for damaged product → Process in accordance with procedure → Send new product
Update product image for Halloween → Adapt existing images with Halloween flair → Update website
etc.
Anyone trained in operations or Lean Six Sigma requires about 3 minutes of knowledge on AI to solve this one. Give them any mystery technology and show them the before and after:
What did the process cost before and what does it cost now?
What happened to throughput and quality?
How often does a human need to approve, review, or repair the output?
What value does the completed task produce?
It doesn’t matter if it’s generative AI or a new method to identify unripe or overripe cranberries based on how far they bounce off a wooden plank, the measurement playbook is fundamentally the same. Implementation might be very different, but the intricacies of calculating ROI for processes are familiar.
What are the special catches for measuring ROI for processes?
Human review is not free. Humans in/on/around the loops are baked into the cost of the process. Don’t fall into the trap of only calculating AI tokens as costs.
Diminishing returns will bite you. Suppose you previously updated a product image once per year and the update historically produced a 5% sales lift. Your newfangled AI process now enables you to update it every day, every hour, or… maybe even every second. … Obviously, we are going to be trillionaires. That’s ludicrous, but some processes are clean with a static return until the job is done (arithmetic wins), while others will shift with the frequency you apply them.
New territory leading to unknown benefits. Some companies who have never had a call center before are firing them up to serve customers. Why? The cost to staff them fell because of AI agents. How much do customers value having a call center to reach out to? They have no idea, but they’re moving forward.
Related to these unknown benefits, here’s one playbook I recommend when tackling benefits that become squishy, either for products or processes.
Performance Indicators (but no Key!)
The measurement may be squishy, but it was often squishy before AI. Why did you update the image once per year rather than twice? How did you justify the old cadence? In many organizations, the answer was a mixture of data, expertise, budget constraints, political jockeying, and mental bandwidth.
AI did not create the measurement problem, but due to its empowering nature across organizations, it did make it harder to ignore.
In an ideal world, every project would have one glorious metric that perfectly represents value and balances risk. In this world, the best metrics have three attributes:
They are linked to the objective.
They are difficult to game.
They are easy to collect.
Unfortunately, metrics often allow you to choose two, sometimes one. Bad metrics that satisfy none of the above are also common, but I count on your wisdom and corporate political savviness to dodge those.
This is why I believe in performance indicators more than I believe in key performance indicators. Lines of code is a terrible KPI for software engineers. Target it and you will produce codebases of hot garbage and drive away talent. … But … if a software team produces zero lines of code for a month, maybe we should chat.
Lines of code is a horrible target and a useful signal. Many solid metrics are like this. Sales calls, token usage, cycle times, experiment counts all can help paint a beautiful picture, and AI can even assist in interpreting whether that picture is Van Gogh’s “The Starry Night” or Goya’s “The Third of May 1808.”
Despite their upbeat attitude, the Cleveland sales team may need assistance
If you have a metric that is closely tied to value, hard to game, and cheap to gather, cherish it. It’s a rare and beautiful thing that will likely be spoiled with the passage of time.
If you do not, use a collection of performance indicators. Look for the confluence of evidence. A compass can point in the right direction even if it can’t tell you how far away the destination is.
Productivity Is Where “ROI of AI” Gets Weird
Now we arrive at the hardest bucket.
Productivity includes individual use of Claude, ChatGPT, Cursor, Codex, Grok Bot, OpenClaw, Hermes, or whatever new product launched while I was writing this sentence. It is easy to measure activity. Who uses the tools? How often? How many tokens? How many tasks? How much do employees rave about how much they like a specific tool?
Our challenge is proving value.
I spoke with a CEO of a mid-sized consumer-facing company who was skeptical about AI productivity gains inside his legal and HR teams. His reasoning was fascinating. He stated that he intentionally hired a limited number of attorneys and HR employees because he did not want those functions doing everything they could theoretically do.
If HR had infinite capacity to identify and enforce every possible infraction, it could destroy the organization. If Legal had infinite capacity to review every contract AND push back on every redline AND perfectly control every minuscule risk, it would grind the company into the ground.
As he put it, “Most attorneys have incredible work ethic and they never run out of things to do. But this is not always a good thing.”
So when someone in one of those functions says, “AI made me much more productive,” the CEO’s reaction may reasonably be: Whoa, whoa, whoa. I do not necessarily want you to be more productive.
This applies far beyond HR and legal. More code can create more useless features and more maintenance. More analysis can create more indecision. More marketing content can create more brand noise. More product ideas can overwhelm the people responsible for choosing among them.
Activity ≠ output.
Output ≠ value.
An hour saved ≠ an hour captured.
A saved hour creates capacity, but it only becomes value if the organization redirects or converts it into something valuable. Your employees can finish tasks 80% faster and fill the time with non-value-add activities that will leave your income statement emotionally unmoved.
You Accelerated Everything Except the Bottleneck
Most work moves through some version of three stages:
Define what should be created.
Build it.
Get feedback and decide what changes.
Historically, building was expensive and slow. Because it was the bottleneck, we often defined the work poorly and refined the definition while building. Essentially, in pre-AI it didn’t matter that we were bad at definition as long as the little gears turned slowly enough that everyone had time to argue, learn, and change direction along the way. We called this “agile,” which it sometimes was.
As any operations professional knows from the fundamental principle of the theory of constraints, you can only move as fast as your bottleneck. AI has changed the constraint. If you can build extremely quickly, then definition becomes the bottleneck. Feedback may become another. Perhaps definition was the real bottleneck all along, but slow execution helped conceal our weaknesses.
This creates what I have been calling an “unequally yoked organization.” One part of the company can sprint while another part cannot decide where to go.
I watched the 2016 Ben Hur with my kids and I was challenged to include a relevant meme
Everyone is moving really fast! But somehow the organization is standing still. You 10x’d productivity and sales are… down.
The employee is not lying and neither is your dashboard. Both are simply measuring the wrong level of ROI of the system.
The question is, “Did making this person faster move the constraint that governs the valuable output of the system?” This means effective measurement of ROI for productivity requires measurable organization units of value and cost.
While the implementation varies according to organization type and size, this may require more changes in structure rather than more AI-enhanced fancy dashboards.
Caveat: Productivity Is the Pipeline
There is a potential danger in the previous argument. You can declare individual productivity impossible to measure and then only focus on revenue and cost, jettisoning all other “fluffy” metrics. That would be a mistake, and one of the principal reasons for it is that productivity enhancements are the pipeline to AI-first transformation.
The most common and legitimate “winning with AI” story is the following:
Someone uses an AI assistant to help perform a task.
They improve the prompt, add context, connect a tool, and discover a repeatable technique.
That technique becomes an AI skill.
The AI skill becomes a shared AI workflow.
The AI workflow becomes a managed process that runs at organizational scale.
Productivity → Skill → Workflow → Process
The initial value may be an employee saving 20 minutes, but the scaled process saves the organization 20,000 hours or wholly transforms the user experience. This means productivity should not be evaluated only at the individual level, but also measure what flows downstream, even if it’s much harder to measure and link each step along the way.
If the organization measures only personal time saved and enhancements to personal deliverables, it will miss the compounding value.
Reporting Productivity of AI
I would not report one AI ROI number to the executive team or board. Instead, I would report a portfolio organized by the three buckets.
For product and process, use the foundation of existing approaches, but add the necessary upgrades and nuances for an AI-first world.
For productivity, look for valuable outcomes and impact by team and organizational units, and be on guard for the inevitable “unequally yoked” chaos that inevitably hits every organization during this transformational time.
Finally, leverage a collage of performance indicators instead of overfitting to a metric that will eventually be gamed into oblivion.
You’ve already tried to prove “my team is 10x faster!”
Now prove the company is going somewhere.










