Government & Public Sector

How to Write an RFP for an AI Project

Most AI RFPs describe a technology. The ones that produce good outcomes describe a problem, a baseline, and how success will be measured. Here is the structure that gets you real answers.

Written by Gurubalan G.T. · · 3 min read

A document icon with a checklist of clean boxes, half checked, representing a well-structured tender document.
A document icon with a checklist of clean boxes, half checked, representing a well-structured tender document.

How to Write an AI RFP That Gets You Honest Answers

Most AI RFPs describe a technology they want. The tenders that actually work describe a problem, a measurable baseline, and how success will be judged — and let vendors propose the technology.

Start with the problem, not the solution

Naming a specific technology in the RFP — "we want an AI agent" — invites every vendor to say yes, regardless of whether an agent is the right answer. Describing the problem instead — "enquiries currently take an average of 42 hours to receive a first response, and we want that under 4 hours" — invites a vendor to tell you honestly whether that needs an agent, a simpler automation, or a process change with no AI at all.

Gartner's warning about "agent washing" — vendors rebranding chatbots and RPA as agents — is precisely the outcome a technology-named RFP invites. (Gartner, June 2025)

What the RFP needs, section by section

A measured baseline, not an estimated one. If you don't know your current enquiry response time, error rate, or processing cost, measure it before the RFP goes out. A vendor cannot propose an honest improvement against a number nobody has actually checked.

A defined evaluation method, decided before proposals arrive. The US government's OMB M-25-22 requires exactly this discipline for federal AI contracts: agencies must retain the ability to evaluate performance on a defined cadence, and evaluation data "should not be accessible to the vendor." (OMB M-25-22) That principle transfers cleanly to any RFP: decide how you'll measure success, and keep the test data out of the vendor's hands until after the test.

A data-use clause, from the first draft. State plainly whether the vendor may use your data to train models sold to other customers. The federal standard — permanently prohibiting this absent explicit consent — is a reasonable default to require in any tender.

A named failure-mode question. Ask every bidder directly: which of your past deployments didn't work, and what happened? The UK's own transparency standard establishes a comparable expectation for algorithmic tools sold into the public sector: suppliers "should be comfortable with this level of transparency." (GOV.UK) A bidder who cannot answer has either not delivered much or won't say — either way, useful information before you commit budget.

An exit clause, priced. What does it cost, in money and elapsed time, to move to a different supplier at the end of the contract? Get an answer in the tender response, not after signature, when your leverage has evaporated.

The clause that filters out weak bidders on its own

Ask bidders to name a comparable deployment that did not meet its goals, and describe what changed afterward. This single question does more filtering than the rest of the document combined: it rewards genuine experience, penalises vendors who have only ever demoed, and is nearly impossible to fake convincingly on the spot.

What to avoid

Avoid specifying implementation details you're not qualified to judge — model architecture, specific frameworks, cloud provider. That is the vendor's job to propose and defend, and locking it in early removes your ability to compare genuinely different approaches to the same problem.

Avoid a scoring weight that rewards the lowest price over the clearest evidence of delivery capability. The UK National Audit Office's finding that supplier arrangements contributed to at least 29 years of delay across five government programmes traces directly to contracts "geared to buying inputs rather than outcomes" — and price-weighted scoring is one of the most common ways that mistake gets built into a tender before it's even issued. (NAO, January 2025)

Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions. Every statistic here is linked to its original published source.

Government & Public SectorRFPprocurementtenderAI projects
Considering a build? Describe the process and we will come back with a scope and a cost range — including if our view is that software is not the right answer. Get a range Message on WhatsApp