Method Solutions AI Workforce Results Websites Blog About Book a call
All articles

How to Evaluate an AI Vendor: 12 Questions

Buyer evaluating an AI vendor against a due diligence checklist of questions

Every AI vendor demos well. The demo is not the product. What separates a system that is still working in a year from one that is quietly switched off is who operates it, how it fails, where your data goes, and what happens when you want to leave. Here are twelve questions that surface all four.

Key takeaways

  • Ask who operates the system after launch and what they do each week. A vendor without a named answer is selling you a build, not a working system.
  • Make the vendor walk you through a failure rather than a success. How a system behaves when it is uncertain tells you more than any accuracy claim.
  • For European deployments, insist on a written data processing agreement, a named sub-processor list, a stated retention period, and explicit confirmation that your data does not train public models.
  • Depth of integration is the difference between an assistant and an employee. If the system cannot write into your calendar, CRM or ticketing tool, staff will keep doing the work manually.
  • A vendor who cannot tell you what their system will not do has either not deployed at scale or is not being straight with you.
  • Settle the exit terms before you sign: data export format, ownership of phone numbers, and notice period. Ask when leaving is cheap, not when it is urgent.

On this page

  1. How to use this list
  2. Questions about who operates it
  3. Questions about failure and escalation
  4. Questions about data, GDPR and residency
  5. Questions about integration depth
  6. Questions about measurement and reporting
  7. Questions about pricing, lock-in and exit
  8. The scorecard, all twelve questions in one table
  9. Red flags that are not on the question list
  10. Frequently asked questions

How to use this list

This is written to be useful whoever you buy from, including if that is not us. We publish it because the most common bad outcome in this market is not a buyer choosing a competitor. It is a buyer choosing badly, having a poor experience, and concluding that the whole category does not work. That outcome is worse for everyone selling honest work.

Two practical notes before the questions. First, ask them live rather than in a written request for proposal. Written answers get polished by someone who is good at writing. Spoken answers reveal whether the person in front of you has actually operated a system in production. Second, the value is in the follow-up. Almost every vendor can answer the headline question. The second and third question in each cluster is where the difference shows.

The format below gives each question, what a strong answer sounds like, and what a weak or evasive one sounds like. The weak answers are not strawmen. They are things buyers hear regularly, and some of them sound reassuring until you look closely.

Questions about who operates it

1. Who operates this system after go-live, and what do they do each week?

Strong: "A named account engineer. Every week they review a sample of interactions, triage the escalations, classify why each one happened, ship configuration changes, and send you a report. Here is last month's report for another client with the identifying details removed."

Weak: "The platform is self-service, you can edit everything yourself in the dashboard." That may be true and it may even be a feature, but it means the operational work is yours. Price that in. If you do not have someone who will spend thirty to sixty minutes a week on it, the system will drift. This is the failure pattern we describe in more detail in why AI pilots never reach production.

2. What happens in the first month after launch, specifically?

Strong: A described sequence with dates. Shadow running, then supervised, then autonomous. Weekly reviews. A named point of contact. An expectation set that the first two weeks will surface defects, because they always do.

Weak: "It just works from day one." No deployment against real customers works perfectly from day one, and a vendor claiming otherwise is either inexperienced or managing your expectations rather than your risk. The tell is the absence of any planned review period.

3. Who is on the team, and how many deployments have they run?

Strong: Concrete answers about team size, roles and the number of live systems. Willingness to say "we are small, here is exactly who would work on yours".

Weak: Vagueness about headcount combined with claims of large scale. Also worth noting: a vendor that has run four deployments well is a better bet than one that has run four hundred badly, so small is not itself a problem. Evasiveness is.

Questions about failure and escalation

4. Walk me through what happens when the system does not know the answer.

This is the most informative question in the list. Insist on a walkthrough, not a description.

Strong: "It detects that it is out of scope, tells the caller plainly that it is handing over, transfers to a named person or creates a task with the full context and a response window, and logs the case for review. If the recipient does not answer within the window, this fallback fires. Here is a recording of it happening."

Weak: "The model is very good, it almost always knows." That is an answer about the happy path to a question about failure, and it usually means no escalation ladder has been designed. A system with no defined way to give up will improvise, and improvisation in front of a customer is how trust is lost permanently.

5. What does the system do when it is uncertain rather than wrong?

Strong: A distinction between confidence in understanding and confidence in the answer, with different behaviour for each. Re-asking once, then escalating. Never guessing on anything that commits the business, such as a price, a date or a promise.

Weak: No distinction at all. If a vendor treats "did not hear that clearly" and "does not know the policy" as the same condition, the handling will be wrong for one of them.

6. What will your system not do?

Strong: A ready list. Intents always routed to a human, such as complaints, clinical advice, legal advice, disputes about money. Industries declined. Claims refused. A vendor who has operated real systems has this list because reality gave it to them.

Weak: "It handles everything." Nobody's does. This answer either means limited production experience or a decision not to volunteer the limitations. Both should lower your confidence. A vendor's willingness to name their boundaries is one of the few signals in this market that is hard to fake.

Questions about data, GDPR and residency

7. Where is our data stored and processed, and who are your sub-processors?

Strong: Named regions, named providers, and a sub-processor list that includes the model and telephony providers. A written data processing agreement under Article 28 of the GDPR available on request. A stated position on international transfers.

Weak: "It is all secure and encrypted." Encryption is a control, not an answer to a residency question. Also weak: an inability to name which model provider sits behind the product. You cannot assess a supply chain you cannot see. Our note on GDPR and AI customer service sets out the specific documents to ask for.

8. Is our data used to train models, ours or anyone else's?

Strong: An unambiguous no for public or shared models, backed by the contractual terms the vendor holds with their own providers. If the vendor does fine-tune on your data for your benefit, they should say so explicitly, explain the isolation, and let you opt out.

Weak: "We may use aggregated and anonymised data to improve our services." That phrase covers a wide range of practices and deserves a follow-up. Ask what specifically is retained, for how long, and whether transcripts of customer conversations are included. In a clinic or a law firm, this is not a theoretical concern.

9. What is the retention policy, and how do we handle a data subject request?

Strong: A stated retention period for recordings, transcripts and derived records, configurable to your policy, with a described process for finding and deleting an individual's data on request. Awareness of the disclosure obligations that apply when a person is interacting with an automated system.

Weak: "We keep everything indefinitely in case you need it." Indefinite retention is a liability, not a feature, and it will not survive contact with a data protection review.

Questions about integration depth

10. What does the system write into, not just read from?

Integration depth is the difference between an assistant and an employee, and it is the most commonly overstated part of any AI proposal.

Strong: A specific list. Reads calendar availability including buffers and staff rules, writes the booking, creates and updates the CRM record, logs the transcript against the contact, sends the confirmation and the reminder, and closes the ticket. Named systems, named directions.

Weak: "It integrates with everything via Zapier." That may be technically true and it is sometimes the right answer for a light workflow, but surface-level connections tend to break quietly and rarely handle the exception cases. The follow-up to ask is: after the system handles an interaction, does any human have to touch any system to complete it? If the answer is yes, you will be running two processes in parallel, and one of them will eventually be declared redundant. The distinction is much the same one we draw between a voice agent and a chatbot.

Ask, too, what happens when an integration fails. A calendar API being briefly unavailable is a normal event. The correct behaviour is to detect it, tell the customer honestly, capture the request and complete it when the system recovers. The incorrect behaviour is to confirm a booking that was never written.

Questions about measurement and reporting

11. How is success measured, what is the baseline, and how often do we see it?

Strong: A named metric tied to a business outcome, a baseline captured before go-live, and a regular report in a consistent format. Coverage, containment, escalation rate with reasons, and the outcome that matters to you, such as appointments booked or first-response time.

Weak: Metrics that only flatter the vendor. "Number of conversations handled" is activity, not outcome. So is "hours saved" when it is calculated from an assumption the vendor chose. Push for a number you already track independently.

If a vendor cannot help you establish a baseline, that is itself a finding. The reason we can quote figures for VEGNA Aesthetic Clinic, where answer rate went to 99% from 62% and no-show rate fell to 10% from 20%, is that the "before" numbers were captured before anything was switched on. The full VEGNA breakdown shows the shape of a report that is actually defensible at renewal.

Questions about pricing, lock-in and exit

12. What does the setup fee cover, what is recurring, and what happens if we leave?

Strong: A clear split. The build fee covers discovery, configuration, integrations, testing and staged rollout. The recurring fee covers monitoring, tuning, reporting and support. Usage costs such as telephony minutes are stated transparently, whether bundled or passed through at cost. On exit: a named export format for transcripts, recordings, contact records and configuration, confirmation that you own the phone numbers, and a notice period in months rather than years.

Weak: A large one-off fee with no recurring component. This sounds cheaper and usually is not, because it means nobody is funded to maintain the system after launch and you will be paying again in six months to fix drift. Also weak: any inability to answer the export question directly. Ask about leaving while leaving is hypothetical. The answer is much harder to obtain once the relationship has soured.

On price levels generally, be suspicious of both extremes. Very cheap usually means a thin configuration layer on a generic platform with no operator behind it. Very expensive should come with something specific attached: bespoke integration work, regulated-industry handling, or genuine custom development. For context on what the components actually cost, we set out the arithmetic in our breakdown of AI receptionist pricing.

The scorecard, all twelve questions in one table

Print this, take it to the call, and mark each answer. A vendor scoring strongly on nine or more is worth progressing. A vendor scoring weakly on any of questions 4, 7 or 12 should be treated with caution regardless of the rest, because those three cover failure behaviour, data governance and exit, which are the three areas where a bad answer is expensive to discover later.

# Question Strong answer sounds like Weak answer sounds like
1Who operates it after launch?A named role, a weekly cadence, a sample report"It is self-service"
2What happens in month one?Shadow, supervised, autonomous, with review points"It works from day one"
3Who is on the team?Specific headcount, roles, number of live systemsVague scale claims, no names
4What happens when it does not know?Detect, disclose, hand over with context, fallback"The model is very good"
5What happens when it is uncertain?Re-ask once, then escalate, never guess a commitmentNo distinction between misheard and unknown
6What will it not do?A ready list of excluded intents and industries"It handles everything"
7Where is our data, and who are the sub-processors?Named regions, named providers, Article 28 agreement"It is encrypted"
8Does our data train models?Explicit no for public models, contract terms cited"Aggregated and anonymised improvements"
9What is the retention policy?Stated period, configurable, deletion process defined"We keep everything"
10What does it write into?Named systems, named write actions, failure behaviour"It integrates with everything"
11How is success measured?Business outcome, pre-launch baseline, regular reportActivity metrics and vendor-chosen assumptions
12What does the fee cover and how do we leave?Split build and operate, export format, number ownershipLarge one-off fee, no exit answer

Red flags that are not on the question list

Some signals appear in how a vendor behaves rather than in what they answer.

The demo is a recording. Ask to call the system live, from your own phone, and try to break it. Interrupt it. Give a date ambiguously. Ask for a person by name. Speak with an accent it will not expect. A vendor confident in their product will encourage this. A vendor who steers you back to the scripted flow is telling you something.

No questions about your process. A vendor who can quote a price before understanding your intent distribution, your booking rules and your escalation contacts is quoting for a template. That is fine if a template genuinely fits, but you should know that is what you are buying.

Pressure on the timeline. Discounting that expires this week is a sales technique, not a commercial reality. It is particularly out of place in a market where the main risk to the buyer is moving too fast on a badly scoped project.

Overreach on the replacement narrative. Be wary of a vendor whose pitch is that the system will replace your team outright. In practice the systems that work take the repetitive, high-volume, out-of-hours load and route the rest to people who now have time to handle it properly. Vendors who oversell total replacement tend to have underdesigned the escalation path, because escalation only matters if you accept that humans are still part of the process. The honest framing is that roles change composition rather than disappear, and that is a subject with a serious public debate around it, some of which is worth listening to before you form a view.

Unwillingness to say no. The best signal available is a vendor who tells you that your process is a poor fit and explains why. Nobody enjoys losing revenue, so a vendor who does it occasionally is demonstrating that their advice is worth something. If every question you ask receives an enthusiastic yes, you are talking to a salesperson rather than an operator.

One last suggestion. Ask for a reference who has been live for at least six months in an industry similar to yours, then ask that reference three questions: what broke in the first month, how quickly it was fixed, and what they still handle manually. Every deployment has a rough first month. A reference who claims otherwise has either forgotten or has been carefully selected. The useful information is in how the vendor behaved when things went wrong, which is the only thing you can be certain will happen at some point.

Frequently asked questions

What is the single most important question to ask an AI vendor?

Who operates this after launch, and what specifically do they do each week. Almost every failed deployment traces back to nobody owning the system once the build was signed off. A vendor with a real answer will name the role, the cadence and the deliverable. A vendor without one will talk about the platform being self-service, which means the work lands on you.

How can I tell if an AI vendor is overpromising?

Ask what the system will not do. A credible vendor has a ready list: intents they always route to a human, industries they decline, promises they refuse to make about accuracy. A vendor who says the system handles everything has either not deployed at scale or is not telling you what happens in the tail of real cases, which is where most of the risk lives.

Does an AI vendor need a data processing agreement under the GDPR?

Yes, if they process personal data on your behalf, which any customer-facing AI system does. You need a written data processing agreement under Article 28, a list of sub-processors including the model providers, a stated retention period, and confirmation of where data is stored and transferred. A vendor who cannot produce these on request is not ready for a European deployment.

How do I avoid vendor lock-in with an AI system?

Ask three exit questions before signing: can I export transcripts, recordings, contact records and configuration in a usable format, do I control the phone numbers and accounts, and what is the notice period. Ownership of the numbers and the customer data matters far more than owning the prompts. Get the export format named in the contract rather than promised verbally.

What should an AI vendor's pricing actually include?

Expect a build fee covering discovery, configuration, integrations and rollout, plus a recurring fee covering monitoring, tuning, reporting and support. Usage costs such as telephony minutes or model tokens should be stated transparently, whether bundled or passed through. Be cautious of a large one-off fee with no ongoing component, because it usually means nobody is funded to maintain the system.

What references should I ask for and what should I ask them?

Ask for a client in a similar industry and of similar size who has been live for at least six months. Then ask the reference three things: what broke in the first month, how long the vendor took to fix it, and what they still handle manually. Those answers are far more informative than a satisfaction score, because every deployment has a rough first month.

Bring these questions to us

We would rather you asked all twelve and chose someone else than skipped them and had a bad experience. If you want straight answers to every one of them, book a call.

Book a call

Sources and further reading