How to evaluate an AI product before adopting it: the questions to ask the vendor
Corporate procurement has decades of experience buying software, and almost none buying software that behaves probabilistically, changes with the supplier's updates and drags along an evolving regulatory apparatus. The result shows in the demos: impeccable presentations, always on the same exampl...
Corporate procurement has decades of experience buying software, and almost none buying software that behaves probabilistically, changes with the supplier's updates and drags along an evolving regulatory apparatus. The result shows in the demos: impeccable presentations, always on the same examples, in front of buyers who do not yet have a questionnaire worthy of the object. From a vendor-agnostic position, with no products to defend or attack, we have distilled the questions that in our clients' projects separate solid suppliers from those selling a roadmap, and it helps to organise them into five areas.
What is inside, really
The first area opens the box. Which part of the product is AI and which is traditional logic? The question sounds naive and is not, because the answer immediately reveals supplier-side AI washing, the legacy product rebranded for one marginal feature. Which models does it rest on, its own or third parties', and in the latter case whose? Does the supplier control versions or inherit the upstream provider's updates, with the behavioural shifts that follow? With what notice does it communicate changes affecting outputs?
A serious supplier answers without reticence, net of details protected by intellectual property; one who opposes trade secrecy to the question "who produces the model that will process my data" is asking for an act of faith, and acts of faith in procurement go by another name.
Where the data goes, and what the system learns
The second area is where mistakes cost most. Is your data used to train or improve models, the supplier's or third parties'? The answer must be demanded in writing in the contract, distinguishing between configuration data, operational data and user feedback, because policies often diverge across the three. Where does the data physically reside, with which sub-processors, with which retention and deletion timelines? How does all this reconcile with GDPR and, for the use cases that require it, with sector constraints?
In this area documentation beats conversation: the DPA, the list of sub-processors, security certifications with their real scope. The salesperson reassures, the contract binds.
The evidence, the regulation, the exit
The third area concerns performance evidence, and the rule descends from what we wrote about benchmarks: the supplier's metrics, obtained on the supplier's cases, count as much as the brochures. The right request is a proof of value on your real cases, with your golden set, your success criteria and a defined duration, and the supplier's willingness to sustain it is itself a data point: whoever has a good product rarely fears the test bench.
The fourth area is regulatory. How does the supplier classify its own system under the AI Act, with what written rationale? What documentation does it provide for the obligations that stay with you as deployer, from user-facing transparency to human oversight? The question should be asked today even of suppliers of non-high-risk systems, because the ability to answer measures the maturity of their governance, and because your systems register needs that information anyway.
The fifth area looks at the end from the beginning: in which formats do your data and configurations come out if the relationship closes, in how much time, at what cost? Do the customisations built on top of the product remain yours? It is the lock-in theme, which deserves the dedicated article we will give it, and which at purchase time boils down to one principle: exit conditions are negotiated when the supplier wants you, never when you want to leave.
The method beyond the questions
The questions work inside a process, and the minimal process has three properties. The same questions to every candidate, with answers written down and kept, because the memory of demos is generous with whoever presented best. A score with weights decided beforehand, for the same reasons seen in use case prioritisation. And the presence at the table, from the first round, of the people who will use the product and the people who oversee its risks, because you already know the cost of involving them later.
In our chain this work lives inside the AI Assessment and the selection that follows, where suppliers get evaluated with the same discipline as use cases: calendly.com/fabiolalli/zerofive, or hello@zerofive.ai. And if a negotiation is already underway, the cheapest test of the month fits in a single request: ask the supplier for a proof of value on your data, and watch what happens to the conversation.