Why AI has changed what data migration can deliver

Oct 08, 2026 | min read
By

Abhay Bagai

The conversations I have with data and technology leaders follow the same arc. They describe a legacy reporting estate that has grown beyond anyone's ability to manage: licensing costs piling up, reports duplicated across departments, shadow processes filling gaps because the official tools do not move fast enough. Then I ask when they last seriously considered doing something about it.

Almost always, they have tried before, and it ran longer, cost more, and delivered less than planned. The reason is not primarily a people problem or a planning problem. It is a discovery problem. Legacy estates accumulate complexity in ways no single team fully understands, and manual analysis is too slow, too partial, and too vulnerable to the biases of stakeholders in the room.

This article sets out how AI changes that, covering the full migration lifecycle from discovery and dependency mapping through to semantic modelling, governance, and activation, and why getting the architecture right at the start determines everything that follows.

The architecture that makes it work

The CI&T approach to data modernisation is structured around six phases: assessment and discovery, refinement and specification, agentic migration, agentic governance, data and AI activation, and change management. What connects all six is a central knowledge base that every AI agent in the programme reads from and writes to throughout the entire lifecycle.

The knowledge base holds structured metadata, technical specifications, semantic context, data lineage, and a technical constitution that governs how agents operate. It is machine-readable rather than a document for people to read, which is what makes the lifecycle genuinely AI-native rather than AI-assisted. Agents consume structured context, and every output they produce enriches the knowledge base for the phases that follow.

This is the architecture within which CI&T FLOW operates. FLOW is CI&T's proprietary AI system, and in a migration context, its application spans the full lifecycle: from automated assessment of the legacy estate through to generating persona-specific training content in the change management phase. The consistency of a single underlying system across every phase is what prevents the knowledge decay that typically afflicts long-running migration programmes, where insights from discovery get lost before they reach the teams doing the engineering work.

Assessment and discovery: replacing assumption with evidence

Discovery is where most migrations are won or lost. Traditional discovery relies on interviews and workshops: stakeholders describe what they use, what matters, and what can be retired. The problem is that self-reported usage is systematically unreliable. People overestimate how often they use reports and underestimate the number of duplicates across departments.

The AI-native approach bypasses that by reading the metadata directly. Rather than asking whether a report is important, agents examine how many users have accessed it, how recently, how often, and whether the underlying data is still being refreshed. The resulting picture is objective in a way that no interview process can replicate.

In a recent programme involving a 20,000-plus report estate, metadata analysis revealed that a large proportion of reports had not been accessed in six months. It also identified 40% more scope than two previous partners had found through conventional methods. Both findings were critical: the first allowed the team to reduce the migration scope on the basis of evidence, and the second meant the full complexity of the programme was understood from the outset rather than discovered mid-delivery.

Three artefacts that unlock parallel delivery

Discovery produces three outputs that drive everything downstream. The knowledge base itself is a data domain map that aligns the estate to business value streams and defines data pod ownership, and a value roadmap that prioritises data products by business impact, feasibility, and time to value. Together, they make it possible to run workstreams in parallel and keep the programme aligned to what the business actually needs, rather than working through a technical checklist in sequence.

Automated dependency mapping and the semantic layer

Understanding what exists in a legacy estate is only the first challenge. Understanding how it is connected is where the real complexity lies. A single business metric, such as weekly sales or customer lifetime value, is rarely stored in a single place. It is calculated from data spanning multiple systems, each owned by a different team and with its own transformation logic.
Automated dependency mapping uses AI agents to trace those connections: for every metric in the legacy estate, the agents identify which source systems feed it, which transformations it passes through, and which downstream reports depend on it.

Missed dependencies are among the most common causes of post-migration failure. A report that appears self-contained in the legacy environment turns out to depend on a data feed that was not included in scope. Manual mapping catches most of these. Automated mapping, running across the full estate simultaneously, catches them all.

Why business logic needs a permanent home

Dependency mapping surfaces a consequential architectural question: where should business logic live? In most legacy systems, the rules for calculating a metric live within the reporting tool itself, whether an older enterprise platform like MicroStrategy or a more modern one like Power BI, rather than in a central location. Switch platforms, and you have to find, validate, and rebuild every one of those definitions from scratch. It is one of the reasons platform migrations are so much harder than they look from the outside.

The right approach, and the one CI&T implements in the Databricks ecosystem, is to extract that logic into a centralised semantic layer. Metric definitions are governed once in Databricks and made available to any consuming application. The organisation is no longer tied to any particular visualisation platform. More importantly, AI agents built on that data work from consistent, governed definitions rather than whatever happens to be embedded in the tool they query. This is what makes Databricks architecturally significant beyond its performance characteristics: it provides the semantic foundation that makes both self-service analytics and AI activation possible from a single governed layer.

Agentic migration: what a fourfold productivity gain actually looks like

The migration phase is where the accumulated intelligence of the knowledge base starts paying off at scale.

Rather than engineers manually analysing each report, understanding its logic, rewriting it for the target platform, and testing the output, CI&T FLOW automates the repeatable elements of that process through chained agent workflows: one agent interprets the legacy report logic, a second generates the DDL scripts and code for the target platform, a third runs quality checks against the output, and a fourth flags exceptions for human review. Each agent reads from and writes back to the knowledge base, so the context built during discovery informs every migration decision.

For the 20,000-plus report programme, CI&T built 19 bespoke AI accelerators tailored to the characteristics of that estate. Development processes that previously took days were completed in hours, a fourfold increase in developer productivity, and the programme delivered in five months what industry benchmarks would estimate at over twelve.

The human-in-the-loop principle applies throughout. Engineers review code and conduct QA after agentic migration, experts validate governance outputs, and data product teams conduct user acceptance testing after AI activation. The agents accelerate, but humans make the consequential decisions.

Governance and lineage: built in, not bolted on

Data governance in a migration context is often treated as a compliance exercise at the end. The AI-native approach treats it as a continuous output of the programme. As the agentic governance phase runs, FLOW generates lineage metadata, catalogues data assets, and performs privacy checks. The outputs feed back into the knowledge base, keeping it current as the estate evolves rather than producing a static snapshot that drifts from reality the moment the programme moves on.

The practical value extends beyond the migration itself. Organisations operating under regulatory data requirements need auditable lineage records that demonstrate where data originated, how it was transformed, and who is accountable for it. Building that capability as a byproduct of the migration, rather than as a separate workstream after the fact, means it is accurate from day one rather than reconstructed after the event.

From parity to capability: self-service analytics and AI activation

The goal of modernisation is not to replicate the legacy estate with modern technology. It is to create capabilities that the legacy platform made impossible. Once business logic is centralised in the Databricks semantic layer and data is governed consistently, two things become straightforward that were previously out of reach: self-service analytics and AI activation.

Self-service analytics:
Databricks' Genie capability allows business users to query the semantic layer in plain language, asking questions and receiving answers without waiting for a data team to build a report. This is not a cosmetic improvement to the user experience. It is a structural change in how quickly an organisation can move from a business question to a data-driven decision.

AI activation: With the semantic layer in place, AI agents work from a single, consistent source of truth rather than pulling from fragmented, contradictory datasets across the organisation. The AI use cases an organisation wants to build on top of its data foundation depend entirely on that foundation being trustworthy. A centralised semantic layer with proper lineage and governance is not a prerequisite that can be addressed later. It is the prerequisite.

Change management as a delivery workstream, not an afterthought

Change management runs throughout the programme, not just at the close. Teams moving from familiar tools to new platforms need continuous support tailored to how different groups actually work, not delivered as generic training at the end of a go-live.

CI&T uses FLOW to generate customised learning content for different user personas, adapting language and examples to each group's specific context. Training materials, onboarding videos, and internal knowledge portals are produced as part of the delivery cadence.

In the 20,000-report programme, a conversational support agent was also built so business users could ask questions directly about the migration's status, the new platform's capabilities, or how to replicate something they had done in the legacy system.

The migration that builds, not replaces

The distinction that matters in this work is the difference between a migration that achieves parity and one that creates capability. Parity means the organisation ends up with the same reports, metrics, and analytical limitations, just running on newer infrastructure. Capability means it emerges with a governed semantic layer, auditable lineage, self-service analytics, and a data foundation its AI agents can actually use.

The organisations that will get the most out of the Databricks ecosystem, and out of every AI investment that follows, are the ones that understand that distinction going in. The six-phase lifecycle, with the knowledge base connecting every step, is designed to deliver the second outcome. The AI does not just accelerate the process. It changes what the process produces.

The work of modernising a data estate is not separate from an organisation's AI ambitions. It is the prerequisite for them. Get the foundation right, and the AI use cases that seemed out of reach become straightforward. The organisations investing in that foundation now are the ones that will be the hardest to catch up with.


Abhay Bagai

Abhay Bagai

Head of Data & AI

 

Want to learn more about how CI&T can help you harness your business's potential? Get in touch!