
The Five Questions to Ask Before You Approve an AI Vendor Quote
Key Takeaways
- The quoted number is usually accurate and usually incomplete. That gap is scope convention rather than dishonesty, which is precisely why asking is your job and not the vendor's
- A complete cost is typically 1.3 to 1.5 times the quoted figure, because six categories are conventionally excluded from implementation quotes across the industry
- The answers matter less than whether the vendor has already thought about the questions. A vendor hearing question three for the first time is telling you something more important than whatever they say next
- Ask “how will we know it works” first. It is the only question whose answer determines whether any of the others can be verified after signature
- Volunteered omissions are the strongest positive signal available. A vendor who tells you unprompted what is not included is showing you how they will behave in month seven, when something is genuinely not included
- The most commonly missing line is year two. Quotes describe construction; systems incur operations, and a proposal that treats launch as completion has mispriced the commitment
- Price the exit at the same time as the entry. Ask what you own and what transfers. The answer separates a capability engagement from an indefinite dependency
- A vendor who answers all five well is worth more than one quoting 20% lower, because the gap between quotes is smaller than the gap between a project that ships and one that does not
Introduction: The Quote Is Not the Problem
A proposal lands. It has a number on it. The number is larger than you hoped and smaller than you feared, and the immediate instinct is to evaluate it by comparing it to another number — a second quote, a budget line, a figure someone mentioned at a conference.
That comparison is close to worthless, and not because vendors are dishonest. In our experience most implementation quotes are accurate about the work they describe. The problem is subtler and more structural: quotes describe construction, and what you are actually buying is a system that has to keep working. Those are different commitments with different cost profiles, and the convention across the industry is to price only the first.
So two quotes that differ by thirty percent may be pricing entirely different scopes, and the cheaper one is frequently the one that omitted more. Comparing the numbers tells you which vendor drew a smaller box around the work. It tells you nothing about which engagement will produce a working system.
What follows is a different approach. Five questions, in a deliberate order, each with an explicit account of what a strong answer sounds like and what a weak one reveals. They take about twenty minutes of a call. They are designed to be asked by someone who is not technical, and they will tell you more about the likely outcome than a line-by-line comparison of two spreadsheets.
What a Quote Conventionally Includes and Excludes
Start with the shape of the omission, because it is remarkably consistent across vendors and worth recognizing on sight.

The left column is what vendors price: discovery, design, build, testing, deployment, and an initial training session. This is real work and the estimates are usually defensible.
The right column is what almost nobody prices, and it is worth understanding why each is excluded, because the reasons are legitimate even when the consequences are expensive.
Evaluation set construction is excluded because it requires your domain experts, not the vendor's engineers. A vendor cannot label your edge cases correctly and should not pretend to. But if nobody scopes it, it does not happen, and a system without an evaluation set cannot be proven to work — which means it cannot be approved, which is how projects stall after a successful build.
Data cleanup is excluded because its size is genuinely unknown until someone looks. This is honest. It is also the largest single source of budget overrun in the category, and the honest version is a discovery phase priced separately rather than an assumption that the data is fine.
Ongoing operations is excluded because it is a different commercial motion — a retainer rather than a project. Budget 20 to 30 percent of build cost annually for monitoring, prompt maintenance, evaluation upkeep, and support.
Model version migration is excluded because it is unpredictable and outside the vendor's control. Underlying models are updated and deprecated on the provider's schedule, not yours. Each meaningful change means re-running your evaluation set and sometimes adjusting prompts. Assume at least one a year.
Your team's internal hours are excluded because they are not the vendor's to quote. They are frequently 20 to 40 percent of the true total, and they are the line that most reliably surprises people, because they never appear on any document.
Security and compliance review is excluded because it depends on your obligations. It is ten times cheaper at design time than at launch, and in regulated sectors it can be disqualifying rather than merely expensive.
None of these are hidden. They are simply outside the conventional boundary of an implementation quote, and the boundary is invisible unless you know to look for it.
The Five Questions
The order is deliberate. Each one constrains what the following answers can mean.

1. How will we know it works?
Ask this first, because it is the only question whose answer determines whether the others can be verified after you sign.
A strong answer names a metric and an artifact. Something structurally like: we will build an evaluation set of roughly sixty real examples labeled by your team, and the acceptance criterion is 90 percent field-level accuracy on extraction with all failures routed to review. The specific numbers matter less than the shape — a measurable quantity, a defined dataset, a threshold agreed in advance.
A weak answer describes a demonstration. “We'll show you the system and you can see how it performs.” This sounds reasonable and is the root of most disputes, because it defers the definition of success until after the money is spent, at which point the two parties discover they meant different things. Every argument about whether an AI system is good enough is downstream of nobody having defined good enough while it was still cheap to do so.
The follow-up that reveals most: who labels the evaluation examples? If the vendor says they will, that is a problem — not because they are being deceptive but because they cannot know your edge cases. The correct answer involves your people and a scoped amount of their time, which means the answer to this question should generate a line item in the next one.
2. What is not included here?
The most informative question in the set, and the one where the delivery matters more than the content.
A strong answer volunteers the omissions against the six categories above, unprompted and specifically. “This does not include data cleanup, which we cannot size until discovery. It does not include ongoing operations, which we would quote separately. And it does not include your team's time, which we estimate at two days a week from a domain expert for the first six weeks.”
A weak answer is “everything is included.” This is never true, and it is diagnostic rather than merely optimistic. It means either the vendor has not thought carefully about the boundary, or they have and are choosing not to raise it. Both predict the same month-seven conversation, where something needs doing, nobody scoped it, and the relationship absorbs the friction.
What you are testing here is not honesty in the abstract. It is how this vendor behaves when the answer is unwelcome — and you would rather discover that now, over a question with no money attached, than later over one with a great deal.
3. What does year two cost?
This question reprices the entire decision, and it is the one most often skipped because the project framing makes launch feel like the end.
A strong answer separates three things: licensing or seats, inference or usage, and operations. It gives an annual operations figure as a percentage of build cost. It notes which parts scale with volume and which are fixed.
A weak answer treats the system as finished at launch. AI workflows are not fire-and-forget. Models change, data drifts, edge cases surface, evaluation sets go stale. A vendor who has not thought about year two has either not operated a system through a model deprecation or is not planning to be there when you do.
The sharpest follow-up in the entire set: what happens to our costs if usage doubles? A vendor who can answer immediately has modeled it. A vendor who cannot has quoted a build without understanding the operating economics of the thing they are building.
4. What happens when the model changes?
A technical question with a commercial answer, and a reliable separator between vendors who have run production AI systems and vendors who have delivered demos.
A strong answer describes a process: we re-run your evaluation set against the new version, compare results, adjust prompts where needed, and deploy after the regression passes. It may include a rough time estimate and note whether it falls inside the operations retainer.
A weak answer has not considered it. Sometimes phrased as reassurance — “newer models are better, so it should improve.” Newer models are usually better on average and are frequently different in ways that break a specific prompt tuned to a specific behavior. Without a regression process, you discover this through a quality complaint from a user weeks after the change.
This question also indirectly re-tests question one. A vendor cannot have a credible regression process without an evaluation set, so an inconsistency between the answers to one and four tells you which was aspirational.
5. Who owns this when you leave?
Ask it plainly. The discomfort is informative.
A strong answer names artifacts and a transfer. You own the prompts, the evaluation set, the integration code, and the documentation. Here is the handover package. Here is what your team would need to be able to do. Vendors who build for capability transfer say this readily because they have done it before.
A weak answer is a retainer. Not because retainers are wrong — many organizations reasonably choose ongoing support — but because a retainer that exists by default rather than by choice is a dependency you did not price. The distinction is whether continuing is your decision or your only option.
The concrete test: ask whether prompts and evaluation sets are delivered in your repository or theirs. It is a small, specific question and the answer is unambiguous.
How to Read the Answers Together
Individually each question is useful. Read as a set, they reveal something the individual answers do not.
Questions one and four are coupled. A regression process presupposes an evaluation set. If the vendor described a robust evaluation approach in question one and had nothing for question four, the first answer was probably assembled for your benefit rather than describing standard practice.
Questions two and five are the same question at different distances. Both ask how the vendor behaves when candor is not commercially convenient. Consistency across them is a stronger signal than the quality of either alone.
Question three is the reality check on the others. A vendor promising evaluation infrastructure, regression processes, and knowledge transfer while quoting no ongoing cost is describing work nobody is being paid to do, which means it will not happen.
The composite you are looking for is straightforward: did this vendor have these answers before I asked? A vendor constructing an answer in real time is not necessarily bad, but they are telling you that this project will be where they develop the practice. That may be an acceptable trade at the right price. It should be a conscious one.
Why the Cheaper Quote Often Costs More
The counterintuitive dynamic worth naming explicitly, because it drives a large share of bad procurement outcomes.
When two vendors quote the same nominal scope and one is meaningfully cheaper, there are only a few explanations. Lower rates. Greater efficiency from having built the same thing before. Or — most commonly — a smaller box drawn around the work.
The third is the default and it is the hardest to detect from the document, because both proposals list similar-sounding phases. The difference is in what each assumes is somebody else's problem. A quote that assumes clean data, no evaluation work, no operations, and no handover will always be cheaper than one that includes them, and it will produce a system that either does not reach production or reaches it and quietly degrades.
This is why the five questions are worth more than a spreadsheet comparison. They surface the boundary, and once both quotes are normalized to the same scope the ranking frequently reverses. The gap between vendors on price is typically 20 to 30 percent. The gap between a project that ships and one that stalls is total.
Three Answers That Should Stop the Conversation
Most weak answers are reasons to probe further. Three are different, and each is a signal to stop and reconsider rather than negotiate.
“We'll use your data to improve the model.” Unless you have explicitly agreed to this with legal review, it should not appear in a proposal. It is not automatically disqualifying, but it requires a documented, deliberate decision rather than a clause nobody discussed.
“Accuracy will be around 95 percent.” Offered before anyone has seen your data, this is a number with no basis. Accuracy is a property of a task, a dataset, and a threshold. A confident figure quoted in advance of any of those is a sales artifact, and it will become the standard you hold them to in a dispute neither side can resolve.
“You won't need to change your process.” Automating a step inside a badly designed process produces a faster badly designed process. Nearly every successful AI workflow involves some redesign of the surrounding steps. A vendor promising none is either not planning to look closely or is planning to raise it after signature.
What to Do With the Answers
Convert them into the document. This is the step that turns a good conversation into a good engagement.
The success criterion from question one becomes an acceptance clause. Written, with the metric, the evaluation set size, and the threshold.
The omissions from question two become either scoped line items or explicit exclusions. Either is fine. Silence is not.
The year-two figure from question three goes to whoever approves budgets, before signature rather than after.
The regression process from question four becomes an operations obligation, with a stated trigger and rough turnaround.
The ownership answer from question five becomes a deliverables list. Prompts, evaluation set, code, documentation, in your repository.
A vendor who answered all five well will have no objection to any of this, because you are writing down what they already told you. Resistance at this stage is itself the answer to a question you did not have to ask.
The Brightter Perspective
We are on the other side of this conversation regularly, and the honest observation is that the buyers who ask these questions are easier to work with, not harder. A scoped engagement with a written acceptance criterion and an explicit list of exclusions is a better project for everyone involved, because the failure modes that damage vendor relationships are almost always ambiguity rather than disagreement.
The pattern we see most often in organizations that had a bad prior experience is not that they chose a bad vendor. It is that nobody established what done meant, so the project ended when the budget did rather than when the outcome arrived. That is a procurement failure more than a delivery failure, and it is preventable in a twenty-minute conversation.
At Brightter, we would rather lose a deal at the scoping stage to a buyer who decided the work was not yet worth doing than win one that stalls at approval because success was never defined. The five questions above are the ones we would want asked of us.
Conclusion
Do not evaluate an AI proposal by comparing its number to another number. The numbers describe different boxes drawn around the same work, and the smaller box is usually the cheaper one.
Ask how you will know it works, and get a metric and an evaluation set. Ask what is not included, and listen for whether the omissions are volunteered. Ask what year two costs, because the system does not finish at launch. Ask what happens when the model changes, because it will. Ask who owns it when they leave, because the answer separates a capability engagement from a dependency.
Then write the answers into the agreement. A vendor who answered well will not mind, and a vendor who minds has answered a question you did not have to ask.
If you are holding a proposal and are unsure what it does not cover, that review is a short conversation and a significantly cheaper one than the alternative. Start a project at brightter.com/start-a-project.



.avif)
































































































