Engineering & Architecture

Data Migration: Where Modernisation Projects Actually Die

The application layer gets the planning attention. The data underneath it is where modernisation projects actually run aground, usually because nobody looked closely at it until the migration was already underway.

Written by Subash R · · 4 min read

A narrow bridge over a chasm with visible cracks partway across, representing the risk concentrated in the data migration phase of modernisation.
A narrow bridge over a chasm with visible cracks partway across, representing the risk concentrated in the data migration phase of modernisation.

Modernisation plans are usually written around the application: the new architecture, the new framework, the new interface. The data migration gets a line item and a few weeks. Then the project reaches that phase and discovers the data was never as clean as the schema implied, and the few weeks becomes the majority of the remaining timeline.

This is a companion piece to our legacy modernisation guide.

Why the data is worse than anyone assumes

A legacy system's data reflects every workaround, every manual correction, and every edge case that was ever handled outside the system rather than inside it. Fields get repurposed for meanings their names no longer describe. Validation that exists in the new schema was often never enforced in the old one, so records exist that the new system will reject on arrival. None of this shows up by reading documentation, because the documentation describes what the system was supposed to do, not what actually happened to the data over years of real use.

The GAO's review of the federal government's oldest legacy systems is instructive here for a different reason than its headline maintenance figures: systems 23 to 60 years old have had that many years for data quality to drift from whatever the original design intended. (GAO-25-107795) A system does not need to be federal or ancient for this pattern to apply — five years of undocumented workarounds produces the same category of problem at a smaller scale.

The specific ways migrations fail

Discovering data quality problems during the migration instead of before it. By the time bad records surface, the team is already under deadline pressure, and the temptation is to write a quick transformation rule to force the data to fit rather than to understand why the exception exists in the first place. Quick transformation rules accumulate the same way the original workarounds did.

Treating the migration as a single event rather than a process with checkpoints. A full cutover, with the old system switched off and the new one live, leaves no way to compare outputs between the two systems once the switch happens. If a subtle transformation error only manifests weeks later, there is no old system left to check against.

Underestimating reconciliation. After data moves, someone has to verify it moved correctly — not by spot-checking a sample, but by reconciling totals, counts and key relationships between old and new. This is unglamorous work that is very easy to under-scope, and it is exactly the kind of task DORA's research suggests even elite teams chronically underinvest in relative to new feature work. (DORA)

Assuming the migration tooling is the hard part. ETL tooling is largely a solved problem. The hard part is deciding what to do with the exceptions the tooling surfaces, and that decision requires business knowledge that often exists only in the heads of people who have been asked to focus on the new system, not the old data.

What reduces the risk

Profile the data before scoping the project, not during it. A data quality audit — null rates, duplicate detection, referential integrity checks, distribution of values against what the schema assumes — belongs in the planning phase, priced as its own line item, because what it finds will materially change the migration timeline.

Run migrations incrementally where the architecture allows it. Migrating one entity type, one business unit, or one geography at a time creates checkpoints where problems are caught at a scale still small enough to fix without threatening the whole programme. This is the same logic behind the strangler fig pattern applied specifically to data.

Keep the old system queryable after cutover, even if it is no longer being written to. The ability to compare a disputed record against its original source, weeks or months after go-live, is worth far more than the cost of keeping a read-only copy running.

Build reconciliation into the plan as a phase with its own sign-off, not an afterthought. Someone specific should be accountable for confirming the data is right, separate from whoever built the migration, because the person who built it is the least likely to notice their own transformation errors.

The evidence on why this matters at the programme level

The largest peer-reviewed study of IT project costs found that cost-overrun risk shows no statistically significant relationship to project size (p = 0.863) — meaning a migration that looks modest in scope is not inherently safer. (JMIS, 2022) Data migration risk in particular does not scale down neatly with project size, because a small system can still carry decades of undocumented data drift.

What we do

We treat data profiling as a distinct, priced phase before committing to a migration timeline, and we build reconciliation into the plan as its own deliverable with its own owner. Most of the migration failures we have seen were visible in the data months before the migration itself began — for anyone willing to look.

If you are planning a migration and want the data risk assessed properly before a timeline is committed, that is a conversation we are glad to have.


Kaizen Spark Tech designs and delivers software, AI, automation and digital infrastructure for businesses and institutions. Every statistic here is linked to its original published source.

Engineering & Architecturedata migrationlegacy systemsmodernisationengineering
Considering a build? Describe the process and we will come back with a scope and a cost range — including if our view is that software is not the right answer. Get a range Message on WhatsApp