Method Solutions AI Workforce Results Websites Blog About Book a call
All articles

How to Measure AI ROI Honestly

AI ROI measurement dashboard showing baseline metrics against post-deployment results

Measuring AI return on investment properly requires one thing most deployments never do: recording a baseline before launch. Everything else, attribution, conversion into money, avoiding double-counting, follows from that. Without it, every number produced afterwards is an estimate wearing the costume of a measurement.

Key takeaways

  • Capture the baseline before deployment: answer rate, response time, enquiry-to-booking conversion, no-show rate, resolution time and cost per interaction. Retrofitting a baseline afterwards is guesswork.
  • "Hours saved" is the weakest common metric because saved time is only worth money when it is redeployed, and it is almost never audited. Measure work that previously did not happen at all.
  • Attribution should be stated, not hidden. Where several things changed at once, report a range and name the confound rather than claiming a precise figure.
  • Apply a recapture discount: some missed enquiries would have come back anyway, so counting all of them as recovered inflates the result.
  • Separate leading indicators (answer rate, response time, escalation rate) from lagging ones (revenue, retention, review volume). Leading indicators move in week one, lagging ones take a quarter.
  • An honest monthly report shows raw counts, states which figures are modelled rather than measured, includes the full cost, and lists what went wrong.

On this page

  1. Why AI ROI measurement usually gets skipped
  2. What to capture in the baseline, before anything is deployed
  3. The metric table: baseline, measurement and how each one gets faked
  4. The attribution problem, and how to handle it honestly
  5. Why "hours saved" is a weak metric
  6. Leading and lagging indicators
  7. How to avoid double-counting
  8. A worked example, using illustrative assumptions
  9. A measured case: VEGNA Aesthetic Clinic
  10. Review cadence and what a monthly report should contain
  11. Frequently asked questions

Why AI ROI measurement usually gets skipped

There is a structural reason most AI vendors do not push hard on measurement, and it is not incompetence. Measurement introduces the possibility of a disappointing answer. A vendor whose commercial model depends on renewal has a quiet incentive to keep the conversation on capability demonstrations rather than outcomes.

Buyers collude in this more often than they admit. Setting up measurement means agreeing in advance what would count as failure, and that is an uncomfortable conversation to have in week one of a project you have just championed internally. It is far more comfortable to deploy, feel that things are better, and describe the improvement in adjectives.

The result is a market where a large share of deployments cannot answer a simple question from a finance director: what did this recover, and how do you know. We think insisting on the answer is the correct position for a vendor to take, and we would rather be measured and occasionally look ordinary than be unmeasurable and always look excellent.

The rest of this note is the method: what to record, how to convert it into money without cheating, and what an honest report looks like.

What to capture in the baseline, before anything is deployed

The baseline window should be at least four weeks and ideally eight, and it must end before any change goes live. If your business is seasonal, note which season it covers, because you will need that context later.

Six measures cover most front-office and support deployments.

Answer rate

The share of inbound contacts that reached a person, expressed as a percentage of total inbound attempts. Take it from the telephony system rather than from memory. The number almost every business quotes from memory is higher than the number in the logs, which is itself a useful finding to surface early. Count abandoned calls and out-of-hours calls, because those are exactly the ones an agent will change.

Response time

Median and ninetieth-percentile time to first meaningful response, per channel. Median tells you the typical experience, the ninetieth percentile tells you the experience of the customers most likely to leave. Speed matters commercially: Harvard Business Review's widely cited analysis of online sales leads found firms responding within an hour were markedly more likely to reach a decision-maker than those that waited longer, a finding we discuss in more detail in our note on the five minute lead response rule.

Conversion from enquiry to booking or sale

Of the contacts that were answered, what share became a booking, a quote accepted, or an order. This is the multiplier that converts recovered contacts into money, so it needs to come from your own data rather than an industry average. If you only have it for one channel, say so.

No-show and cancellation rate

Bookings that did not turn up, as a share of bookings made, split by lead time if you can. This is one of the cleanest metrics available because it is unambiguous, sits in your calendar system already, and responds quickly to reminder sequences. We go into the mechanism in reducing no-shows with automated reminders.

Resolution time and first-contact resolution

For support-heavy businesses: median time from ticket open to resolved, and the share resolved without a second contact. First-contact resolution is the more honest of the two, because resolution time can be improved simply by closing tickets faster than customers are actually satisfied.

Cost per interaction

Total loaded cost of the function divided by interactions handled. Loaded means salary plus employer contributions plus software plus the share of overhead you would normally allocate, not just the headline wage. Most businesses have never calculated this and are mildly surprised by the result. It matters because it is the denominator in the comparison that comes later, and because it is the figure a finance director will check first. Our breakdown of what an AI receptionist actually costs sets out the equivalent cost side for the AI system.

The metric table: baseline, measurement and how each one gets faked

The final column is the most useful part of this table. Every metric here has a well-worn way of being made to look better than it is, and knowing the trick is the fastest way to read a vendor report critically.

Metric How to baseline it How to measure it after Common way it gets faked
Answer rate Telephony logs over four to eight weeks, counting all inbound attempts including out-of-hours and abandoned. Same source, same definition, equal-length window, same seasonal position where possible. Silently redefining the denominator to exclude out-of-hours or repeat calls, so the rate rises without any behaviour changing.
Response time Median and ninetieth percentile per channel, from system timestamps rather than staff recollection. Same percentiles, same channels, reported together. Reporting the mean, or reporting only the median, which hides the long tail where customers are actually lost.
Enquiry to booking conversion Bookings divided by answered enquiries, from the CRM or calendar, over the same baseline window. Identical calculation. Track new enquiries separately from existing customers. Counting existing-customer calls as new enquiries, which inflates the volume of "recovered leads" considerably.
No-show rate No-shows divided by booked appointments, from the calendar system, split by lead time. Same, over an equal window. Watch for changes in booking mix. Counting a rescheduled appointment as attended rather than as a no-show avoided, which credits the system twice.
Ticket resolution time Median open-to-resolved, plus first-contact resolution share. Same two figures, always reported as a pair. Closing tickets aggressively to cut the median while reopen rate quietly climbs.
Cost per interaction Fully loaded cost of the function divided by interactions handled. Full AI cost, including build amortisation, retainer and usage, divided by interactions. Comparing the AI subscription price against a human salary while omitting build, integration and management time.
Revenue recovered Not baselined directly. Derived from the four measures above. Recovered contacts multiplied by measured conversion multiplied by average value, then discounted for recapture. Applying no recapture discount, so every previously missed call is counted as a customer who would otherwise have vanished.
Hours saved Time study or reasonable estimate of hours spent on the function. Only meaningful if you can name where the hours went instead. Multiplying estimated hours by an hourly rate and presenting the product as cash, with no redeployment evidence.

The attribution problem, and how to handle it honestly

Attribution is the hardest part, and the part where most reporting quietly falls apart. Businesses rarely change one thing at a time. An AI agent goes live in the same quarter that a new website launches, ad spend rises, a competitor closes, or the season turns. Any of those could move the same numbers.

There are four defensible responses, in descending order of rigour.

Run a holdback. Leave part of the system unchanged for the first month: one location, one phone line, one channel, or one time window such as weekends only. You then have an internal comparison group experiencing the same market conditions. This is the closest a small business gets to a controlled test, and it costs almost nothing. The objection is always that it delays benefit on the held-back slice, which is true and usually worth it.

Use year-over-year comparison rather than month-over-month. Comparing October to the previous October removes most seasonal distortion, which month-over-month comparison does not. It requires that you have clean historical data, which is a good reason to fix your reporting before you deploy anything.

Isolate the directly traceable subset. Some outcomes are unambiguously attributable: a booking created by the agent, at 21:40, from a caller who had never contacted you before, on a line that previously went to voicemail after hours. That booking is not a matter of interpretation. Report the traceable subset separately from the modelled total, and let the reader weigh them differently.

State the confound and give a range. When all else fails, honesty outperforms precision. "Bookings rose 18%. We also increased ad spend by 30% in the same period. The share directly traceable to after-hours calls answered by the agent is 7%, and we would attribute somewhere between 7% and 12% to the deployment." A finance director will trust that report far more than a clean-looking single number, because it shows the author was looking for the alternative explanation rather than avoiding it.

Why "hours saved" is a weak metric

Hours saved is the default AI ROI metric because it is easy to produce and flattering. It is also close to meaningless in a small business, for three reasons.

First, it is almost always self-reported and almost never audited. "This saves us about ten hours a week" is an impression, not a measurement, and impressions of time are notoriously unreliable.

Second, saved time only becomes money if it is redeployed into revenue-producing or cost-reducing work. In a five-person business, an hour freed from answering the phone frequently disperses into general slack: slightly longer breaks, slightly less pressure, a marginally calmer day. Those are real quality-of-life gains and worth having, but they are not cash, and presenting them as cash is exactly the sort of thing that destroys credibility when someone eventually checks the bank balance.

Third, multiplying hours by an hourly rate assumes the cost was variable. If nobody's hours were reduced and nobody was not hired, the wage bill is identical. The saving is notional.

Three stronger alternatives:

  • Work that previously did not happen at all. Calls answered after hours, quotes chased on day three, review requests sent. This is the strongest category because the counterfactual is clean: the volume was zero.
  • Rate changes on existing work. Conversion, no-show rate, first-contact resolution. These are ratios from your own systems and are hard to argue with.
  • Cost per interaction. A single figure that lets you compare an AI system, a human team, an answering service and any mixture of the three on the same basis.

If you do want to report time, report it as capacity released and name the destination. "Reception recovered roughly six hours a week, which went into pre-treatment consultations" is a claim someone can check. "Saves ten hours a week" is not.

Leading and lagging indicators

These move on different timescales, and conflating them causes both premature celebration and premature panic.

Leading indicators respond within days: answer rate, response time, contacts handled, escalation rate, containment rate (the share of interactions the agent completed without a handoff). They tell you whether the system is working mechanically. They do not tell you whether it is worth money.

Lagging indicators take a quarter or more: revenue, retention, customer lifetime value, review volume and rating, and the eventual reduction or non-increase in headcount cost. These tell you whether it was worth money, but they are slow and noisy.

The practical approach is to hold the leading indicators to a hard standard immediately and give the lagging ones ninety days before drawing conclusions. If the leading indicators are not where you expected them by week three, do not wait for the lagging ones to confirm bad news. Fix the system.

One warning: it is possible to have excellent leading indicators and no commercial result. An agent can answer 99% of calls beautifully and produce nothing if the extra calls answered are suppliers and wrong numbers. That is why the enquiry-type split in the baseline matters.

How to avoid double-counting

Double-counting is the most common arithmetic error in AI ROI reporting, and it usually happens innocently, by adding up benefits that overlap.

Three specific traps.

Counting the same booking twice. If the agent answered a call that became a booking, and the reminder sequence then prevented that booking from becoming a no-show, you have one appointment, not two. Count the booking once at full value and count the no-show prevention only on appointments that would have been booked anyway.

Counting recovered revenue and saved labour cost together. If the case for the deployment was that staff time was freed, and the same staff then converted more enquiries, some of the recovered revenue is the product of the freed time. Adding both at full value inflates the total. Pick the primary mechanism and treat the other as secondary.

Ignoring the recapture rate. Not every missed call is a lost customer. Some callers ring back. Some leave a voicemail that gets returned. Counting 100% of previously missed calls as recovered revenue overstates the result, sometimes by a wide margin. If you have no data on your recapture rate, apply a conservative assumption, state it explicitly, and try to measure it properly in the next cycle.

A worked example, using illustrative assumptions

The figures below are illustrative and modelled, not measured client results. They exist to show the shape of the calculation. Substitute your own numbers wherever a figure appears.

Assume a clinic with the following baseline, recorded over eight weeks before deployment:

  • 400 inbound calls per month
  • 70% answered, so 120 calls per month unanswered
  • Of answered calls, 60% are new enquiries and 40% are existing customers, suppliers or wrong numbers
  • Measured conversion from answered new enquiry to booking: 25%
  • Average value of a first appointment: EUR 180

After deployment, the answer rate reaches 98%, so 112 additional calls per month are answered.

The calculation, step by step:

  1. Additional calls answered: 112.
  2. Apply the enquiry-type split. 60% of 112 is roughly 67 genuine new enquiries. The other 45 are real calls but not revenue events.
  3. Apply the measured conversion rate. 25% of 67 is roughly 17 additional bookings.
  4. Apply a recapture discount. Assume 35% of those callers would have reached the clinic eventually by calling back. That leaves 65%, or roughly 11 genuinely incremental bookings.
  5. Convert to revenue. 11 bookings at EUR 180 is roughly EUR 1,980 per month in incremental revenue.
  6. Subtract the full cost. Assume a retainer of EUR 850 per month plus EUR 140 of usage, giving EUR 990.
  7. Net monthly contribution: roughly EUR 990, a ratio of about two to one on the money spent.

Two observations about this example. The recapture discount in step four removed a third of the apparent benefit, and it is the step most commonly omitted. Without it the same deployment would report roughly EUR 3,060 of recovered revenue and a three to one ratio, and the number would be wrong.

The second observation is that the answer is sensitive to average transaction value, not to the AI. Run the same model with an average value of EUR 60 and the deployment barely covers its cost. Run it at EUR 400 and it is transformative. This is why a vendor quoting a universal ROI multiple should be treated with suspicion: the multiple is a property of your business, not of the software. The same logic applies when comparing against hiring, which we work through in AI front office versus hiring a receptionist.

A measured case: VEGNA Aesthetic Clinic

Against the illustrative model above, here is a measured deployment. VEGNA Aesthetic Clinic in Amsterdam had a baseline answer rate of 62%. After deployment, 99% of calls were answered. The no-show rate moved from 20% to 10%. The measured recovered revenue figure was EUR 6,050 per month.

What makes those numbers usable is the same thing that makes them unremarkable to produce: the 62% and the 20% were recorded before anything changed. The clinic could see its own answer rate in its telephony logs and its own no-show rate in its calendar. Nothing had to be reconstructed after the fact or estimated from an industry benchmark. The full detail sits on the VEGNA results page.

This is the entire argument of this article in one example. The difficulty of proving value is almost never analytical. It is administrative, and it is decided in the fortnight before go-live, when someone either does or does not export the current numbers.

Review cadence and what a monthly report should contain

Measurement without a cadence decays into an annual argument. Three loops work well.

Weekly

Leading indicators only, and it should take ten minutes: answer rate, response time, escalation rate, containment rate, plus any interaction where the customer abandoned. This is an operational check, not a financial one.

Monthly

The full comparison against baseline, with a decision attached. Widen scope, narrow scope, or leave unchanged. A monthly review that never changes anything is a status update.

Quarterly

Lagging indicators, and a genuine re-examination of assumptions. Is the recapture rate you assumed still plausible now that you have more data. Has the average transaction value shifted. Would you buy this again at this price.

What belongs in the monthly report

  • The baseline figures, restated in full every month. Not linked, restated. People forget them.
  • Raw counts behind every percentage. "98%" means nothing without "392 of 400".
  • Every assumption used to convert activity into money, labelled as an assumption.
  • A clear separation between measured figures and modelled figures.
  • Full cost, including the retainer, usage, and any internal time spent managing the system.
  • Escalation rate and its trend.
  • A section on what went wrong: failures, complaints, awkward interactions, anything the agent handled badly.

That last item is the tell. A monthly report with no failures in it is not a report, it is marketing. Every system in production has bad interactions in any given month, and a partner who does not surface them either is not looking or has decided you should not see them. Neither is a good sign, and the businesses that get the most out of these systems are the ones that read the failure section first.

Frequently asked questions

How do you measure the ROI of an AI system?

Record a baseline before deployment, then compare the same metrics over an equal-length period afterwards. Capture answer rate, response time, conversion from enquiry to booking, no-show rate, resolution time and cost per interaction. Convert the change into money using values you already hold, such as average first-visit value, then subtract the full cost of running the system. Without a baseline recorded before launch, any figure you produce afterwards is an estimate dressed as a measurement.

Why is hours saved a weak metric for AI ROI?

Because saved hours are only worth money if they are redeployed into something that produces revenue, and in small businesses they usually get absorbed into general slack. Hours saved is also self-reported and almost never audited. Better measures are the ones that show up in a system anyway: revenue recovered from work that previously did not happen, conversion rate on enquiries, no-show rate and cost per interaction.

How long should you run an AI system before judging the ROI?

Allow roughly two weeks of tuning that you exclude from measurement, then measure a full ninety days. Two weeks is too short because early behaviour is unrepresentative and seasonality dominates. Ninety days covers a full billing cycle, at least one quiet period and one busy one, and is long enough for lagging indicators such as retention and review volume to start moving.

How do you handle attribution when several things changed at once?

State the confound rather than hiding it. If you also changed pricing or increased ad spend in the same period, the honest report says the combined effect and shows which parts are directly traceable. Where possible use a holdback: leave one location, one channel or one time window unchanged for the first month so you have an internal comparison group. A clearly stated range beats a precise number nobody trusts.

What is a realistic ROI for an AI receptionist or front office?

It depends almost entirely on call volume and average transaction value, so any vendor quoting a universal multiple is guessing. The arithmetic is straightforward: recovered enquiries multiplied by your measured booking rate multiplied by your average value, discounted for customers who would have called back anyway, minus the full monthly cost. Businesses with high call volume and high transaction value see the strongest returns.

What should an honest monthly AI report contain?

The baseline alongside the current period, the raw counts behind every rate, the assumptions used to convert activity into money, the full cost including the retainer, the escalation rate, and a section on what went wrong. A report with no failures listed is a marketing document. It should also state clearly which figures are measured and which are modelled.

Get your baseline recorded before you deploy anything

We will pull your current answer rate, response time and no-show rate first, so whatever happens next is measurable rather than arguable.

Book a call

Sources and further reading