Engineering & Architecture

Rewrite, Refactor, or Replace?

These three words get used interchangeably in planning meetings and they describe very different risk profiles. Confusing them is how a two-month refactor quietly becomes an eighteen-month rewrite.

Written by Subash R · · 4 min read

Three paths diverging from a single starting point toward three different-shaped destinations, representing the rewrite, refactor and replace decision.
Three paths diverging from a single starting point toward three different-shaped destinations, representing the rewrite, refactor and replace decision.

These words get used interchangeably in planning conversations, and the imprecision is expensive, because the three options carry genuinely different risk profiles. This is a companion piece to our broader legacy modernisation guide — here we focus narrowly on choosing between the three most commonly confused options.

What each one actually means

Refactor changes the internal structure of code without changing what it does from the outside. Behaviour is preserved; the improvement is in maintainability, testability and the speed of future changes. This is the option with the smallest blast radius, because a correct refactor is, by definition, invisible to the system's users.

Rewrite discards the existing implementation and builds a new one intended to do the same job, usually with the same or expanded scope. This is a full re-creation of business logic that may be poorly documented, encoded only in the current system's behaviour, and understood by fewer people than anyone assumes.

Replace discards the system entirely and adopts a different solution — often a commercial product — rather than building an equivalent. This converts a build problem into a buy problem, with its own trade-offs, covered in our build vs buy framework.

Why teams default to rewrite when refactor would do

The honest answer is that refactoring old, poorly-tested code is harder and less satisfying than writing new code. Refactoring requires understanding what the existing system actually does, including behaviour nobody intended but that some downstream process now depends on. A rewrite lets you skip that understanding and start fresh — which feels productive right up until the rewritten system quietly breaks something the old one handled correctly by accident.

This is a real and well-documented failure mode, sometimes summarised as "the second-system effect": teams rewriting a system tend to over-scope it, adding capabilities the original never had, because building something new invites ambition in a way that maintaining something old does not.

The evidence on project risk applies to all three, unevenly

The largest peer-reviewed study of IT project costs — 5,392 projects collected, 4,677 with usable cost data — found no statistically significant difference in cost-overrun risk by project size (p = 0.863). The largest overrun recorded was a $1,500 project that finished at $425,000. (JMIS, 2022)

That finding matters more for rewrites and replacements than for refactors, because a refactor's scope is naturally bounded by the existing system's actual behaviour — there is a defined target to preserve. A rewrite's scope is bounded only by what the team decides to include, which is exactly the kind of ambiguity the JMIS data associates with the long tail of bad overruns.

A decision test that works most of the time

Ask what the actual problem is. If the problem is "the code is hard to change safely," that is a refactoring problem, and refactoring solves it directly without re-litigating decades of business logic. If the problem is "the system cannot do something the business now needs, and no amount of internal restructuring will make it able to," that is a genuine case for a rewrite or a replacement — but confirm that conclusion by trying to name the specific architectural limitation, not just the general frustration with the codebase's age.

If you cannot name the specific limitation, you probably have a refactoring problem wearing a rewrite's ambition.

What to check before committing to a rewrite

Is the current behaviour documented anywhere other than the code itself? If not, budget significant time for behavioural discovery before writing a line of the replacement, because the risk is not building the new system — it is building the wrong new system because nobody wrote down what the old one actually did.

Can the two systems run in parallel during transition? A rewrite with no parallel-running period concentrates risk into a single cutover moment. Our piece on the strangler fig pattern covers the alternative.

Who signs off that the new system is behaviourally equivalent, and how? Without an answer to this, "done" becomes a subjective judgment made under deadline pressure, which is how regressions reach production.

What we do

We default to refactoring wherever the underlying problem allows it, because it is the option with the lowest risk of quietly losing correct behaviour nobody remembered to specify. When a rewrite really is warranted, we treat behavioural discovery as its own phase, priced and scheduled, rather than folding it invisibly into development time.

If you are trying to work out which of these three your system actually needs, that is a conversation we are glad to have.


Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions. Every statistic here is linked to its original published source.

Engineering & Architecturelegacy systemsrefactoringtechnical debtengineering
Considering a build? Describe the process and we will come back with a scope and a cost range — including if our view is that software is not the right answer. Get a range Message on WhatsApp