The short answer
Fifteen questions sort the AI services market: three on production proof, three on the team, three on process, three on ownership, three on integrity. The pattern in the answers matters more than any single one: operators respond with specifics, dates, and named trade-offs; resellers of demos respond with adjectives. Ask in order and take notes.
Vendor conversations are asymmetric: they have answered buyer questions a hundred times, you are asking them for the first time. This list is the equalizer, grouped so the conversation flows naturally, with the good and bad answers spelled out. It applies to everyone, including us.
Group 1 · Production proof
- 01Show me a system of yours running in production right now. Good: a URL you can use today, or a live walkthrough in a real environment. Bad: a demo video and a case study PDF. Demos are the easiest artifact in AI to manufacture.
- 02What broke in the last quarter, and how did you find out? Good: a specific incident, how it was detected, what changed after. Operators have war stories with dates. Bad: nothing ever breaks. That means no monitoring or no production.
- 03What do your evals look like? Good: golden datasets, thresholds, re-runs gating every change, and an example of evals catching a regression. Bad: the subject changes, or evals are an optional line item.
Group 2 · The team
- 01Who exactly does the work, and where do they sit? Good: names, seniority, and time inside your workflow. Bad: a bench description, or discovery conducted entirely through a brief. Distance from the work is the leading documented cause of failure.
- 02Which role shows up after signature? Good: the same engineers you met, embedded. Bad: pre-sales engineers through the deal, a ticket queue afterwards. The role distinction has its own field guide.
- 03How many engagements run at once per engineer? Good: a number that makes embedding plausible. Bad: evasion, which prices in divided attention you were not told about.
Group 3 · The process
- 01How does discovery happen? Good: inside the workflow, with the people who run it, producing a prototype. Bad: workshops and a paid analysis phase that never touches the actual queue.
- 02What gates a release to real users? Good: eval thresholds, a defined first cohort, monitoring switched on before launch. Bad: the client signs off on a demo. That is a feeling, not a gate.
- 03What happens when the model provider retires a version? Good: pinned versions, deprecation dates in the diary, a rehearsed migration with golden re-runs. Bad: a blank look. Every LLM system inherits this event on a schedule.
Group 4 · Ownership
- 01Walk me through the handover, item by item. Good: code, prompts, infrastructure, evals, credentials, runbooks, named individually, with a cold-start test at the end. Bad: hesitation on any item, especially prompts and evals.
- 02Where does the system run, and whose accounts hold the keys? Good: your cloud, your repositories, your API keys, from the start or transferred whole. Bad: hosting that must stay with the vendor. That is rent with extra steps.
- 03What happens if we part ways mid-engagement? Good: a defined exit with work-to-date transferred. Bad: ambiguity, which you will discover at the worst possible moment. Who owns what, legally, deserves its own read.
Group 5 · Integrity
- 01When did you last tell a prospect not to use AI? Good: a recent, specific story where the honest answer was process change or boring automation. Bad: AI fits every problem you describe. Some of what you described is a cron job.
- 02What is this NOT good at yet? Good: crisp limits, and design implications for each. Bad: capabilities all the way down. Vendors fluent in limits have operated systems; the rest have operated slideware.
- 03What would make you refuse this project? Good: missing workflow owner, unreachable data, volume too low to repay the build. Bad: nothing would. A vendor with no refusal conditions has no scoping discipline.
Operators answer with specifics, dates, and named trade-offs. Resellers of demos answer with adjectives.
How to run the conversation
Ask in order: production proof first, because a weak answer there makes the rest theater. Take notes on specifics, not vibes; you are counting dates, names, and numbers. Compare vendors on the same fifteen, and weigh the pattern rather than any single stumble. The deeper context for each group lives in the partner-selection field guide, the disqualifiers are compressed in the red flags list, and the ownership group has a checkable standard in the ownership handover. We publish our own answers across this site, which is the standard we think the question list should hold everyone to.