Business & Enterprise

How to Choose a Software Development Partner

The statistic everyone quotes about outsourcing failure rates comes from a 2003 Gartner prediction with no published methodology. Here is what the peer-reviewed evidence actually says about picking a technology partner.

Written by Gurubalan G.T. · · 7 min read

Three solid geometric shapes, with a magnifying glass revealing the middle one to be hollow, illustrating vendor due diligence.
Three solid geometric shapes, with a magnifying glass revealing the middle one to be hollow, illustrating vendor due diligence.

You have probably read that 50 to 60% of outsourcing deals fail.

That figure is usually attributed to a Gartner prediction from around 2003, always reported secondhand, always without a published methodology. We could not trace it to a primary source. Neither could we trace "70% of outsourcing failures are due to poor planning," which is widely attributed to Deloitte and appears nowhere in any Deloitte report we could find.

This matters more than pedantry. If you are about to commit a serious budget, the numbers shaping your caution should be real. Most of the numbers in this category are not.

Here is what survives scrutiny.

What the evidence actually supports

The most rigorous work on this question is a systematic review by Lacity, Khan, Yan and Willcocks, which examined 164 empirical articles published between 1992 and 2010 across 50 journals, coding 741 tested relationships. Its broad conclusion is sobering for anyone hoping for a formula: outcomes depend far more on how the relationship is structured and governed than on any attribute of the vendor. (Journal of Information Technology, 2010)

The finding that has held up best comes from the earlier Lacity and Willcocks study of 61 sourcing decisions across 40 US and UK organisations, drawing on 145 participants: selective outsourcing outperformed both total outsourcing and total insourcing on achieving expected cost savings, and decisions made jointly by senior executives and IT managers outperformed unilateral decisions by either group alone. (MIS Quarterly, 1998)

Read that again, because it inverts how most selection processes run. The failure mode is not choosing the wrong vendor. It is choosing the wrong shape — handing over everything, or nothing — and doing it without the people who understand the technical reality in the room.

There is also a signal in the AI-specific data. Project NANDA's 2025 study of enterprise AI implementation found that external partnerships reached deployment around 67% of the time, compared with about 33% for internally built tools, with employee usage rates nearly double for externally built systems. Handle with care: it is a preliminary working paper rather than peer-reviewed research, the interview sample is 52 organisations, the figures are self-reported, its authors build AI infrastructure, and the report itself notes that correlation does not prove causation.

What nobody can tell you is which specialist. That part is on you, and it comes down to the questions you ask.

The four questions that reveal the most

1. "Which of your projects didn't work, and what happened?"

This is the single most informative question you can ask, and almost nobody asks it.

Anyone who has delivered at scale has a project that underdelivered. They can tell you what went wrong, at what point they knew, what they did about it, and what they changed afterwards. That answer demonstrates three things simultaneously: real delivery history, a functioning post-mortem culture, and the confidence to be candid with a prospect.

Someone who cannot answer has either not delivered much, or does not want to say. Both are worth knowing before you sign.

Watch for the evasive version: a failure story where the client was entirely at fault. That is not candour, it is blame-shifting with better staging.

2. "What would make this cost three times your estimate, and what catches it early?"

The largest peer-reviewed study of IT project costs — 5,392 projects collected, 4,677 with usable cost data, completed between 2002 and 2014 — found that the median project lands on budget (overrun ratio 1.0) while the mean is 1.8, because overruns follow a power-law distribution with a very long tail. (JMIS, 2022)

The study also tested whether smaller projects are safer. They are not: differences in the tails across project sizes were not statistically significant (p = 0.863). The largest overrun in the dataset was a $1,500 workflow customisation that finished at $425,000.

So the risk is real, unpredictable by size, and the only defence that generalises is early detection. A partner who has internalised this will answer immediately and specifically — naming the integration nobody has tested, the data quality nobody has checked, the approval nobody has scheduled. A partner who says "we have a rigorous process" has not answered.

3. "What does it cost me to replace you?"

Switching cost is real and almost never modelled at the point of purchase.

The UK Competition and Markets Authority found that in cloud services, less than 1% of customers switch provider each year, with technical and commercial barriers that "lock customers into their initial choice of provider." (CMA, July 2025)

The same dynamic applies to development partners, minus the regulator. Ask specifically: Who owns the code, in writing? Where do the repositories live and who has access? Is the documentation good enough for a competent stranger to pick up? What is the transition support, and is it priced?

The answer you want is unhesitating, because a partner who has thought about your exit has thought about your interests.

4. "Who will actually do the work?"

The most common bait-and-switch in this industry is not fake credentials. It is the senior team who wins the work and the junior team who delivers it.

Ask for the names, the seniority mix, and how much of their time you are getting. Get it in the contract. Ask what happens if a key person leaves.

What a proposal should contain

A quote that states a single number and a timeline is a position, not an estimate. A proposal worth taking seriously shows:

Assumptions, stated openly. Team composition, seniority, duration, rate. You cannot challenge what you cannot see, and a partner confident in their reasoning will show it.

A quality assurance line. Software QA analysts and testers earn a mean of $111,490 annually in US BLS data — cheaper than developers, not free. (BLS) A proposal with no QA is not cheaper, it is incomplete.

A realistic view of capacity. DORA's research found that even elite engineering teams report spending only about 50% of their time on new work, the rest going to unplanned work, rework and support. (DORA) A plan assuming full-time feature delivery is a plan that will slip.

The running cost. Hosting, dependency updates, support, and the change the business will inevitably need. If year one is the only year priced, you are looking at a deposit.

A discovery phase with an honest description of what it does. Worth noting: the claim that early investment prevents defects costing "100 times more" later rests on a citation to an IBM institute that was an internal training programme and published no dataset. The largest modern study — 171 projects between 2006 and 2014 — found "no evidence for the delayed issue effect." (Empirical Software Engineering, 2017) Discovery is still worth doing, because requirements quality correlates with outcomes in the studied contexts. But a partner selling it with the 100× claim is selling you folklore.

Red flags

A price with no workings. See above.

No failure story. See above.

Certainty about your requirements in the first meeting. Anyone who fully understands your problem before investigating it is pattern-matching to something they have already built.

Reluctance about transparency obligations. If you are a public body, the UK's transparency standard states that suppliers selling algorithmic solutions to public bodies "should be comfortable with this level of transparency." Hesitation is informative.

Pressure on timeline. Discounts that expire are a sales technique, not a commercial reality.

Claims about AI capability that don't survive a follow-up question. Gartner estimates that of the thousands of vendors claiming agentic AI capability, only around 130 are real, and names the practice of rebranding chatbots and RPA as agents. (Gartner) Ask what the system does when it is uncertain. Real agent architectures have an answer.

A note on how we would want to be judged

We have written this knowing you might apply it to us. That is the point.

We answer the failure question, because we have projects that taught us things. We state our assumptions so they can be challenged. We price the running cost alongside the build. And we write the exit into the start — your data, your repositories, your documentation, in portable formats, on request, without a commercial conversation attached.

If you are shortlisting partners and want a conversation that starts with your problem rather than our capabilities, we are glad to have it.

Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions. Every statistic here is linked to its original published source, with its sample size and limitations stated where they matter. Where a widely-quoted figure turned out to be untraceable — including two in the opening paragraph — we have said so rather than repeated it. If you find something here we have got wrong, tell us and we will correct it in public.

Business & Enterprisevendor selectionoutsourcingdue diligenceprocurement
Considering a build? Describe the process and we will come back with a scope and a cost range — including if our view is that software is not the right answer. Get a range Message on WhatsApp