Six months into an AI subscription, someone in the room says “it’s definitely saving us time.” Everybody nods. Nobody has a number. The renewal gets approved because canceling feels like going backward, and that is how a tool nobody uses turns into a permanent line item.

You do not need a spreadsheet model or a consultant to fix this. You need four small measurements and the discipline to take the first one before you start. Here is how to tell whether your AI tools are actually earning their keep, and how to avoid the metrics that will happily tell you what you want to hear.

Start With One Task and a Stopwatch

The single biggest mistake is measuring “AI” as a category. AI is not a thing you can measure. A task is. Pick one job that happens often, that you can describe in a sentence, and that somebody does the same way every time. Writing a proposal. Producing the weekly service report. Summarizing a discovery call. Sorting the shared inbox each morning.

Then time it before you change anything. Have the person doing it record the minutes, five or ten times, over a normal couple of weeks. This is the step everyone skips and the only one that makes the rest meaningful. Once people start using a new tool, their memory of the old way gets generous. A baseline you wrote down beats a baseline you remember.

After the tool has been in use for a month, time the same task again, the same way. Multiply the difference by how often the task happens and by a loaded hourly cost for that role. That number, per task, is the closest thing to honest ROI a small business is going to get. It is also usually enough to make the decision obvious in either direction.

Count the Rework, Not Just the Minutes

Speed with a bad output is not a savings, it is a transfer. The work moves from the person doing it to the person fixing it, and sometimes that second person is more expensive than the first. So track a second number alongside the clock.

  • How often did the output need real correction? Not a comma. A factual error, a wrong number, a paragraph that had to be rewritten. Count the times, not the severity, and keep it simple.
  • How long did checking take? Review time is part of the task now. If a draft arrives in two minutes and takes twenty five to verify, you did not save eighteen minutes off a thirty minute job.
  • Did anything reach a customer wrong? This gets its own line. One wrong quote or misstated commitment can wipe out a quarter of time savings, and it is the risk your reputation cares about.
  • Is the correction rate falling? Early on, high correction is normal while people learn to give better instructions. If it is still high after six weeks, the tool is not a fit for that task.

This is the practical version of treating AI like a junior staff member. You expect to check a new hire’s work, and you expect the checking to shrink over time. If it never shrinks, that is information you only get by counting.

Track Seats Used Against Seats Paid For

Here is the least glamorous and most reliable measurement available to you. Per seat software gets bought for the whole team and used by a third of it. You are paying for all of them.

The good news is that vendors give you this data. Microsoft’s admin center documentation, published in 2026, describes a Copilot usage report with exactly the metrics you want: “enabled users,” meaning unique users with a license over the selected period, “active users,” meaning enabled users who actually tried a Copilot feature, and an “active users rate,” which Microsoft defines in that 2026 documentation as active users divided by enabled users. The report covers 7, 28, 90 or 180 day windows and breaks usage down by application.

That active users rate is your utilization number, and it is the fastest way to find money on the table. Microsoft’s adoption guidance, also published in 2026, recommends reviewing trends, noting users or time periods with low or declining usage, and notifying inactive licensed users who might need guidance, with its example targeting people who did not use the product within a rolling thirty day window. Other vendors publish similar admin reporting, so check what yours offers.

When you find inactive seats you have two honest options and one dishonest one. Train those people, because plenty of nonuse is confusion rather than rejection. Or reclaim the licenses and cut the bill. What you should not do is leave it alone and keep calling the rollout a success. Our guide on preparing your team for AI without the hype covers the training half.

The Metrics That Will Fool You

Some numbers feel like progress and measure nothing. They are easy to collect, which is exactly why they show up in status updates.

  • Number of prompts. Microsoft’s 2026 documentation does include total prompts submitted and average prompts per user, which help spot who has stopped entirely. As a success measure they are backwards. Ten prompts to get one usable paragraph is worse than one prompt, not better.
  • Logins. Opening a tool is not doing work with it. Attendance is not output.
  • Words or documents generated. Volume of output is only good news if somebody needed that output. Otherwise you have automated the production of things nobody reads.
  • Enthusiasm. The two people who love any new tool are not your sample. Ask the skeptics, because they are the ones who will tell you it did not help.
  • Vendor supplied time savings estimates. Useful for building a hypothesis, worthless as evidence about your business. Only your own before and after numbers count.

The Two Honest Questions at Renewal

When the invoice comes around, put the dashboards aside and ask two questions in a room with the people who actually use the tool.

First: would we notice if this disappeared tomorrow? Not “would we miss it.” Would work visibly slow down, would something stop getting done, would a customer wait longer? If the honest answer is that Tuesday would look about the same, you have your verdict regardless of what the usage report says.

Second: would the team fight to keep it? People argue to keep tools that save them from work they hate. Silence in that conversation is a very clear answer. And if only three of your twelve users speak up, that is not a failure, that is a signal to buy three seats instead of twelve.

Both questions get at the same thing our post on why saving time takes priority over saving money argues: the point was never the software. The point was hours returned to people who had better things to do with them.

The Bottom Line

Measuring AI ROI is not complicated, it is just easy to avoid. Pick one task. Time it before. Time it after. Count the rework. Compare seats used to seats paid for using reporting your vendor already provides. Ignore prompt counts and enthusiasm. At renewal, ask whether you would notice its absence and whether anyone would fight for it. Do that honestly and you will pay for the tools that work and drop the ones that do not, which is the whole goal. Confirm your vendor’s current reporting features and pricing first, since both change often.

We help businesses across Denton County set up this kind of measurement before the rollout instead of after the invoice, so the decision at renewal is based on evidence rather than a feeling. If you would like help figuring out whether your current tools are pulling their weight, we are glad to help. Contact us today


Sources:

Comments are closed

This website uses cookies and asks your personal data to enhance your browsing experience. We are committed to protecting your privacy and ensuring your data is handled in compliance with the General Data Protection Regulation (GDPR).