The short answer

The demo is designed to work.

Every AI demo you will ever see runs on data the vendor chose. That isn't dishonest, it's how demos work. The problem is that your buying decision depends entirely on the cases the demo left out: the malformed documents, the ambiguous ones, the fifteen percent where the right answer requires knowing something about your business that isn't in the file.

Knowing how to evaluate AI vendors comes down to testing that gap before you sign. I take your real historical cases, including the ugly ones, and I make the vendor run them. I ask how output quality is measured and ask to see the numbers. I price the run cost at your volume rather than the volume in their pricing page. And I read the contract for the clauses that decide what happens when the system is confidently wrong, which is the scenario nobody negotiates until it happens.

I take no referral fees and no vendor relationships, which is the only reason this advice is worth anything. Sometimes the recommendation is to buy from the cheapest one on the list. Occasionally it's that all four should be declined and the problem solved another way. I serve US and Canadian companies, and I've sat on both sides of this table: I've bought these systems and I've built them.

Also known as: AI vendor evaluation, AI due diligence checklist, AI vendor selection, AI procurement support, AI RFP evaluation, AI platform assessment.

What goes wrong

Four ways a good-looking
vendor disappoints.

None of these are visible in a sales cycle. All of them are findable in a two-week evaluation.

The thin wrapper

A prompt, a model API and a nice interface, priced like proprietary technology. Legitimate as a product, badly overpriced as one.

Accuracy with no denominator

"Ninety-four percent accurate." On what test set, scored by whom, and what happens in the other six percent?

The pilot that can't scale

It works at a hundred documents a day. At four thousand the latency, the cost curve or the human review queue makes it unusable.

Your data, their moat

The contract lets them train on your inputs and gives you no export path. Two years in, switching costs more than the original build.

The diligence

Five areas, and the
questions that separate them.

The full checklist goes into the report. These are the areas it covers and the questions that do the most work.

01

Capability, tested

Whether it works on your cases, not on theirs.

  • Run fifty of our real historical cases, including the failures
  • How do you measure output quality, and can we see the test set?
  • What does the system do when it isn't confident?
  • Show us a customer at our volume and let us call them
02

Architecture & substance

What is actually theirs, and what is a model API with a subscription attached.

  • Which model providers do you depend on, and what happens if one changes pricing?
  • What have you built that a competent team couldn't rebuild in a quarter?
  • How do you version and evaluate changes to prompts and models?
  • Where does our data go, and which subprocessors touch it?
03

Economics at your volume

The number that decides whether you keep this in year two.

  • Total cost per transaction at our actual monthly volume
  • What happens to the price when volume doubles, and when it halves?
  • How much human review does the workflow still require?
  • What's the all-in cost of the integration work on our side?
04

Risk, security & compliance

The part legal will ask about after the build is done, so ask it first.

  • Data residency, retention and whether inputs are used for training
  • Security posture, penetration testing and incident history
  • How the system's decisions are logged for audit
  • Alignment with the obligations you already carry to your own customers
See the governance work
05

Contract & exit

Negotiated before signature, because afterward you have no leverage at all.

  • Performance commitments with a remedy attached, not just an SLA on uptime
  • Export of your data and your configuration in a usable format
  • Price protection at renewal and on volume growth
  • Who is liable when a wrong output causes a real loss
The engagement

Two weeks,
before you sign.

Longer for a platform decision with four or more vendors in scope.

Days 1–3

Frame the decision

  • Agree what the system has to do and how you'll know it did
  • Assemble fifty real cases, weighted toward the hard ones
  • Set the comparison criteria before anyone sees another demo
  • Establish the run-cost model at your real volume
Days 4–8

Put the vendors through it

  • Technical sessions with each vendor's engineers, not the account team
  • Run the test cases and score the results the same way for everyone
  • Reference calls with customers at comparable scale
  • Security and data-handling review
Days 9–12

Score and negotiate

  • A single comparison table with evidence behind each cell
  • Contract review focused on performance, exit and liability
  • Negotiation support, including the questions that move price
  • A written recommendation with the reasoning and the risks

Make every vendor run the same fifty of your ugliest historical cases. The scoreboard writes itself, and it rarely matches the demo.

Oshri Cohen
Common questions

About evaluating vendors.

How do I evaluate AI vendors and AI companies?

How to evaluate AI companies selling into this space, in one sentence: test them on your data before you compare their pricing. Assemble fifty real historical cases weighted toward the difficult ones, make every vendor run the same set, and score the results yourself against criteria you set before the demos started. Then price the run cost at your actual volume, ask how they measure output quality and to see the test set, and call a reference customer at comparable scale. The vendor who welcomes that process is usually the one worth buying from.

What questions should I ask an AI vendor?

Five do most of the work. What is your accuracy measured against, and who assembled the test set? Where does a case go when the system isn't confident? Total cost per transaction at our volume, including the human review we'd still be doing? Which model providers are you dependent on, and what happens when one of them changes pricing? And if we leave, can we export our data and configuration in a usable format? Vague answers to the last two are the most reliable warning sign in the whole process.

Is an AI automation agency legit?

Plenty are. The category also has a low barrier to entry, so it attracts people whose entire capability is assembling a low-code workflow. The distinction is not whether they build on top of an existing model, since almost everyone does and there's nothing wrong with it. The distinction is whether they can tell you how quality is measured, what it costs to run at your volume, and who owns and debugs the system after launch. An agency with an evaluation process and a named handover plan is legitimate. One that judges quality by whether customers complain is selling you a demo.

What belongs on an AI due diligence checklist?

Five areas. Capability tested on your own cases rather than the vendor's. Architectural substance, meaning what they actually built versus what they subscribe to. Economics at your volume, including the human review the workflow still needs. Risk, covering data residency, training rights, logging and security. And the contract terms for performance remedies, data export and liability when an output causes a loss. The report I produce works through all five with evidence rather than assertions.

Do you take referral fees from vendors?

No, and I don't hold vendor partnerships. It's the only structure under which this advice is worth paying for. I'm regularly the person telling a client that the expensive option is right, and just as regularly the person telling them that none of the four should be signed and the problem is better solved another way.

What does a vendor evaluation cost?

It's quoted to scope and depends on how many vendors are in the running and how complex the decision is. A two-vendor comparison is usually one to two weeks of work; a platform decision with four vendors and a security review runs closer to three. Against a contract in the six figures, the cost of the evaluation is small next to the cost of picking wrong. That is the argument, rather than anything above it.

Got quotes you
can't compare?

Send me the shortlist and what you're trying to achieve. I'll tell you which questions would separate them fastest, whether or not you hire me to ask them.

hello@oshricohen.me(514) 777-3883Fort Lauderdale · Montreal