A proposal is in front of you. It has a price, a timeline, a diagram of the architecture, and a vendor who answered every question in the meeting well. What you lack is any way to tell that vendor apart from one who is equally good at meetings and will not deliver. Both produce the same artifact, and the artifact is where you are asked to decide.
Three tests separate them, and none of the three depends on your ability to evaluate the technology. They are about the form of what the vendor hands back, not its content, which is what makes them usable by a buyer who is not an engineer. Ask them in this order:
- What runs at the end, and against which of our data?
- What is the pass condition, who wrote it, and when was it written?
- Where does the code live, and what happens if you stop?
Each question asks for something that can fail in front of you, which no proposal can.

Why A Document Is The Wrong Thing To Judge
A document has no failure condition, and nothing can grade it before the money moves.
Research on AI project failures (RAND, 2024) locates the damage not in the technology, but in the gap between the business outcome someone approved and the technical system that was reachable in the buyer's environment. A document cannot reach that gap. Enterprise survey data (BCG, 2024) found that across roughly 1,000 executives, only about a quarter of companies had moved past proof of concept into anything of value.
Test One: Something Running In Your Environment
The first test is whether what the vendor hands you can be wrong in front of you. A diagram cannot. A system running end to end against your named data, in your environment, can.
Here is that difference in a real system. We built the cloud backend for HeartSciences' AI-ECG platform, which had to integrate in both directions with a major EHR vendor. Two hospitals on the same vendor sent the patient identifier in the same field — one zero-padded the medical record number to a fixed width, the other left it variable — and a parser that worked at the first site mismatched results to the wrong patient at the second. One site expected results only in answer to an explicit order; the next, unsolicited.
None of that was in a specification or would have appeared in any architecture document, however good. It appeared because code ran — against a full EHR integration simulator we built so the team could keep shipping, not queuing for vendor sandbox windows.
Two answers should end the conversation. If the vendor names an artifact rather than an execution — a report, a blueprint, a set of recommendations — you are buying a description of work. If the vendor will only run against sample data they supply themselves, the one environment that decides the outcome has been excluded from the test.
Test Two: Targets Written Before Anyone Starts
The second test is about timing, and it is the one most vendors fail quietly. A pass condition written after the work is a grade the vendor gave itself. A pass condition written before the work, agreed by both sides, is the only version that can produce an uncomfortable answer.
Ask for the targets in writing before anything begins: the workflow the system must handle, the data it must handle it against, the path it must traverse, and the conditions under which you would say this did not work. A stranger who did not help write those sentences should be able to read them and grade the result. If not, it is not a target.
Two commercial terms follow. The first is the price, agreed in writing before work starts; a range that survives the engagement means it did not reduce the risk it was hired to reduce.
The second is when you pay. Build work runs as a monthly engineering partnership, billed in advance, with targets agreed before the release started: you watch the system meet them and run your own checks in your own test setup. Interim results early rather than a finished document late is what the Quantitative Researcher at the World Gold Council named in his published reference: "Came back fast with interim results and nailed the delivery."
Test Three: Where The Code Lives And Who Runs It
Buyers most often postpone the third test, which is how a first release becomes a permanent dependency. Ask where the code will live, and when it becomes yours. The correct answer names one repository holding the work from day one, which transfers to you monthly as it is paid for, with its full history, documentation and intellectual property.
Then ask the uncomfortable version — what happens if you stop. If your own engineers, or a different vendor, could take the repository and keep the system running, ownership transferred. If the honest answer requires the original vendor to stay, the transfer did not happen and the price you paid was for access rather than for a system.
Running the system should be a separate decision, not a condition of the build. Ours is $10,000 a month and hands-off: monitoring, incident triage, a standing review of cost and drift, and an intact audit trail. It writes no code and ends on 30 days' notice from either side.
When These Tests Do Not Apply Yet
If several candidate workflows remain with no owner and no written success condition, testing a vendor is early. The work of choosing the first one comes first and costs much less.
The second disqualifier is the data. If nobody can produce it — not a production copy, not a de-identified extract, not a synthetic one with the same shape — your blocker is the data path, not the vendor. That is what the data foundation a build actually needs exists to address, and whether your data is build-ready tells you which of the two problems you have.
Sign Against Evidence, Not Confidence
Something running closes the gap between a description and your environment. Targets written first close the gap between the vendor's judgment and yours. A repository that becomes yours closes the gap between buying a system and renting one.
A proposal that cannot fail also cannot protect you.
To run the three tests on us, bring the workflow you have in mind to a fifteen-minute call.
References
- RAND Corporation. Analysis of AI Project Failures. RAND Corporation, 2024.
- Boston Consulting Group. AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value. BCG, 2024.



