What to Ask an AI Development Partner Before You Sign
The questions to ask an AI development partner before you sign — and what a good answer sounds like. Every AI development company now claims the same capabilities, so the interview matters far more than the pitch deck, and the useful questions are not the ones on the average RFP.
Below is what we would ask if we were buying rather than selling. For each question there is the reason it matters, the shape of a good answer, and the answer that should end the conversation. Several of them are questions we have had to answer ourselves, and one or two are uncomfortable.

Questions About the Work They Have Actually Done
What have you put into production that runs without someone watching it?
Nearly every firm can demonstrate something impressive under supervision. The gap between a demo and a system that runs unattended — recovering from partial failure, handling the input nobody anticipated, leaving an audit trail — is where most AI projects quietly die.
A good answer names a specific system, says what it does on its own, and volunteers what it is not allowed to do without a human. What should worry you is a demo video, a pilot that never left the pilot stage, or a list of models they have "worked with".
Can we see the write-up rather than the logo?
A client logo proves a contract existed, not that the work was any good. A written account of what was built, what was reused and what had to be made from scratch is much harder to fake, and it tells you how the firm thinks.
A good answer is a document you can read without a sales call. What should worry you is a portfolio of logos with no detail behind them, or metrics with no source — a vendor who invents numbers for a past client will invent them about you.
The AI-Specific Questions Most Buyers Skip
These are the ones that separate a team that has run AI in production from a team that has integrated an API. If you ask nothing else, ask these.
Show me your evaluation set.
This is the single most revealing question on the list. An evaluation set is the collection of real cases, with known-correct answers, that a team runs the system against every time they change it. Without one, "it works well" means "it worked the last few times we tried it", and nobody can tell whether today's change made the system better or worse.
A good answer describes how they build one from your data, how big it needs to be before it means anything, and how they handle the cases where reasonable people disagree about the right answer. What should worry you is any version of "we test it manually" or "the model is very accurate" with no way to demonstrate it.
Which parts of this are the model, and which parts are ordinary code?
Good AI engineering uses the model for the part that genuinely needs judgement and plain deterministic code for everything else. Routing, validation, permissions, calculations and business rules should not be left to a language model — they are cheaper, faster and more reliable as code, and they can be tested.
A good answer is a partner who has clearly thought about minimising how much the model is responsible for. What should worry you is enthusiasm for putting the model in charge of the whole workflow, which is expensive, non-deterministic and very hard to debug at three in the morning.
What happens when it is wrong?
It will be wrong. The question is what the system does about it: whether it knows it is uncertain, whether it stops and asks a person, what the fallback path is, and who finds out. This matters more as the system takes on more autonomy — we go into the architecture of that in our guide to agentic AI in production.
A good answer covers confidence thresholds, human approval for the consequential actions, and a log you can audit afterwards. What should worry you is "we use a better model for that" as a complete answer.
What will inference cost at our volume in month twelve?
Build quotes are for building. The bill that surprises people arrives later, from usage, and it scales with how the system was designed — how much context gets sent, how many model calls each task makes, whether results are cached. A design decision made in week two can multiply the running cost for years.
A good answer includes a per-transaction estimate, states its assumptions, and explains which design choices were made to keep it down. What should worry you is a partner who has never had to think about it, which usually means they have never run one of these at volume.
Data, and Who Owns What
Where does our data go, and will it train anyone's model?
The answer depends on which providers sit in the pipeline and on which tier of service is being used. It is a question of fact, not opinion, and any partner who is genuinely running these systems will know it precisely.
A good answer names every provider that will see your data, states the retention position for each, and knows what changes if you have data-residency obligations. What should worry you is reassurance without specifics.
Who owns the prompts, the fine-tunes and the embeddings?
Most contracts settle who owns the code and stop there. AI systems have assets that are not code and are often more valuable: the prompt library that took months to get right, any fine-tuned model weights, the embeddings built from your own documents, and the evaluation set itself. If the contract is silent, you may find you cannot take the working part of the system anywhere else.
A good answer is that all of it is yours and it is written down. What should worry you is a partner who has not considered the question, or one whose platform quietly retains the thing that makes the system work.
Commercials and the People Doing the Work
Will you take a paid discovery phase, and will you fix the price afterwards rather than before?
A fixed price on a vague scope is not the protection it appears to be. It transfers risk to the party with the least information, and the usual result is a partner defending the scope rather than solving the problem. A short paid discovery — process mapping, data assessment, a written recommendation — lets both sides price the real work.
A good answer is yes to both, with a discovery output you keep whether or not you continue. What should worry you is a firm quoting a confident fixed price for something nobody has examined yet. The same logic applies to platform decisions, which we cover in custom CRM/ERP or customise the platform.
Who is actually going to write this?
The people in the pitch are not always the people on the project. Ask for names, ask what else those people are on, and ask what happens if the lead engineer leaves halfway through.
A good answer introduces the delivery team early and is straightforward about how much of their week you get. What should worry you is a senior presence that evaporates after signature.
How many hours of our working day do you actually overlap?
Offshore delivery is not a problem in itself — plenty of enterprise AI work is delivered that way. Undisclosed overlap is a problem, because it decides how quickly a blocked decision gets unblocked.
We will answer this one about ourselves, since it makes the point better than an abstract example. We work Monday to Friday, 10 AM to 7 PM Pakistan time. That is comfortable overlap with the Gulf, a good half-day with the UK, and roughly one hour with the east coast of the United States. We say so before a contract rather than after, and where a US client needs more, that means someone working evenings and being paid for it.
A good answer is a number and a plan. What should worry you is "we are flexible".
What does handover look like if we bring this in-house?
Three Answers That Should End the Conversation
How We Answer These
You should ask this even if you have no intention of doing it, because the answer tells you how the system is being built. Software designed to be handed over is documented, has its credentials and infrastructure defined as code, and does not depend on one person's memory.

A good answer treats handover as a deliverable with a date. What should worry you is anything that sounds like the system will only ever run while they run it.
We publish what we can check and decline to publish what we cannot. Our Hyves case study describes a B2B trading platform we built end to end, including which parts were reusable and which had to be built new, and it carries no invented metrics — the one figure on the page is the client's own published claim and is labelled as theirs. Our approach to building against systems you already run, rather than replacing them, is set out on our enterprise AI integration page.
If you are running this process now, put these questions to us. Ask the evaluation-set one first — it is the one that sorts the field fastest, and we would rather be measured on it than on a deck.
Explore Our Services
Ready to turn these ideas into working software? Our engineering team can help.
AI Agent Development
Autonomous AI agents that automate workflows and scale your operations.
Learn more →Generative AI Development
Custom LLM copilots, content engines and RAG systems built for production.
Learn more →AI Enterprise Integration
AI that plugs into your existing ERP, CRM and procurement stack — integration, not rip-and-replace.
Learn more →All AI & Software Services
Explore our full range of AI, web, cloud and custom software engineering.
Learn more →Start a Project
Book a free discovery call and scope the highest-ROI build for your business.
Learn more →


