Four steps to modernize legacy code with AI, left to right: understand the system, protect its behaviour with characterization tests, decide per part whether to retain, refactor, rebuild or retire, and migrate in small reviewed slices. The subtitle notes that blind AI translation was correct for at most 47.3% of samples.

How to modernize legacy code with AI without breaking business logic

The old system still earns money, and the rules that make it earn live in code nobody fully remembers. AI can help you move it forward, as long as it reads first and converts last.

How do you modernize legacy code with AI without breaking business logic?#

Use AI to read the system and pin its current behaviour with tests first, then change it in small reviewed slices, because blind AI translation was correct for at most 47.3% of samples. That figure comes from Pan et al., "Lost in Translation", an ICSE 2024 paper last revised on 16 January 2024. The authors concluded that "LLMs are yet to be reliably used to automate code translation".

Share of AI code translations that were correctfewer than half at best

2.1%

Lowest-scoring model studied

47.3%

Highest-scoring model studied

Even the best model studied got fewer than half of its translations right.

Show data table
Share of AI code translations that were correct (percent of translations correct)
Optionpercent of translations correct
Lowest-scoring model studied2.1%
Highest-scoring model studied47.3%

Source: Source: Pan et al., ICSE 2024 (arXiv 2308.03109), 16 January 2024

So the question is not whether AI can help, because it can, and a lot. Instead, the real question is where you point it. In practice, AI is strong at reading code, summing it up and repeating a known change many times. However, it cannot tell whether an odd line in an old function is a bug or a rule your finance team depends on.

That gives a four-step order. First, understand the system. Second, protect its current behaviour with tests. Then decide, part by part, what to retain, refactor, rebuild or retire. Finally, migrate in small slices that a software engineer reviews and can roll back. Then each section below takes one step and shows where AI speeds it up and where it must not decide.

How much faster does AI make work on a mature codebase?#

In a randomised trial on mature codebases, developers expected AI to cut task time by 24% and it made them 19% slower. The trial was run by METR and published on arXiv, with version 2 dated 25 July 2025. Sixteen experienced developers worked on 246 tasks in projects they had known for an average of 5 years.

Show data table
Everyone expected a speed-up. The stopwatch measured a slowdown. Source: METR (arXiv 2507.09089), 25 July 2025.
Item Change in task time (%)
Economics experts predicted: faster by 39%
ML experts predicted: faster by 38%
Developers forecast before: faster by 24%
Developers estimated after: faster by 20%
Measured in the trial: slower by 19%

Every forecast was a speed-up of 20% to 39%, and the trial measured a 19% slowdown.

Figure Everyone expected a speed-up. The stopwatch measured a slowdown. Source: METR (arXiv 2507.09089), 25 July 2025. METR (arXiv 2507.09089), 25 July 2025

The gap between feeling and fact is the part to take into a budget meeting. After the 2025 trial ended, the same developers still believed AI had saved them 20% of their time. Also, experts in economics and machine learning had predicted savings of 39% and 38%. In the paper's own words, "AI tooling slowed developers down."

This does not mean AI is useless on old code. Instead, it means a plan built on felt speed is built on the wrong number. Because a legacy system is exactly the mature, familiar codebase the trial studied, the finding applies closely. Therefore, set the budget on what a small pilot slice actually delivers. Treat any vendor multiplier as a claim to test, not a line in the plan.

Why does converting legacy code with AI break business logic?#

A study of 1,700 code samples found 15 categories of translation bugs, and a converter cannot tell which quirk is a bug and which is a business rule. In that 2024 study, Pan et al. translated code between C, C++, Go, Java and Python. The samples came from three benchmarks and two real projects, and even with tidy, well-known code, most translations failed.

DiagramHow a blind conversion handles an undocumented business rule: it copies it or drops it, and nothing records which. Mechanism described from Pan et al., ICSE 2024 (arXiv 2308.03109), 16 January 2024.

However, legacy business code is harder than a benchmark. The rules that matter tend to live in edge cases. Think of a rounding step for one region, a date check added after an audit, or a status code a partner still sends. Nobody wrote them down, because at the time they were obvious. As a result, a model that converts line by line either copies the quirk or quietly drops it. Either way, it does not know which it did, and neither do you.

The same problem shows up in smaller tasks. For example, a MITRE team asked models to add comments to legacy code in their study on documenting MUMPS and mainframe assembly. With a naive prompt, "the models would often alter the code to generate documentation". A tool asked only to explain the code changed it instead.

So AI cannot modernize legacy code on its own. Conversion is not a syntax problem but a behaviour problem. As a result, you need evidence of what the code does today before anything can be trusted to change it.

What should you understand about a legacy system before you change it?#

Before any change, recover four things from the code itself: the business rules, the dependencies, the integrations and the edge cases the data actually hits. The rules are the checks and sums the business depends on. The dependencies are the libraries, frameworks and shared code each part needs. Then come the integrations: every file drop, API call, queue and shared database.

The four things an assessment recovers, what AI drafts for each, and what a software engineer checks it against. Source: method from this article; AI drafting limits from Diggs et al., MITRE (arXiv 2411.14971), accessed 29 September 2026.

What to recoverWhat it isWhat AI can draftChecked against
Business rulesThe checks and sums the business depends onA plain-words explanation of a strange functionProduction logs, real data samples and the people who use the screens
DependenciesThe libraries, frameworks and shared code each part needsA map of which modules call whichThe running system
IntegrationsEvery file drop, API call, queue and shared databaseA list of the endpoints a legacy API exposesProduction logs and the running system
Edge casesThe cases the data actually hitsInputs that reach each branchReal data samples

This is where AI helps most, as a fast reader. For example, it can explain a strange function in plain words, draft a map of which modules call which, and list the endpoints a legacy API exposes. The same MITRE study found model-written comments on legacy code "generally hallucination-free, complete, readable, and useful" next to human ones. However, the authors also found that no automated metric reliably predicted comment quality. In short, a person still has to check the output.

That check is the real work. A software engineer tests each AI draft against the running system, using production logs, real data samples and the people who use the screens. For instance, take a draft that says "orders over a limit need approval". It is only useful once someone finds the limit, the exception for key accounts and the report that reads both.

Google's experience report on AI code migrations, published on 12 January 2025, points the same way. It says AI can "reduce barriers to get started" on a migration, while developers still own the work. As a result, this step ends in a short assessment of what the system does, what it touches and where the risky rules sit.

How do characterization tests protect working behaviour?#

A characterization test records what the code does today, right or wrong, so any change that alters behaviour fails a test before it reaches a customer. It does not ask what the code should do. Instead, it feeds real inputs through the current system and captures the outputs. Then those outputs become the expected result.

  1. Feed real inputs through the current system.

    It does not ask what the code should do.

  2. Capture the outputs.

    Right or wrong, they are what the code does today.

  3. Make those outputs the expected result.

    A known wrong answer is safer than an unknown one.

  4. A later change alters an output and the test fails.

    The change is caught before it reaches a customer.

  5. A person decides whether the change is a fix or a break.

    Without that test, you learn about it from a support ticket.

That sounds odd at first, because some of those outputs may be wrong. However, a known wrong answer is safer than an unknown one. If a later change alters it, the test fails, and then a person decides whether the change is a fix or a break. Without that test, the same change reaches a customer first, and you learn about it from a support ticket.

AI speeds up this step more than almost any other, because writing hundreds of plain tests is slow and repetitive. It can read a function, suggest inputs that reach each branch and draft the test code. Research is moving here too. For example, UnitTenX, published on 6 October 2025, is an open-source system built to "generate unit tests for legacy code". Its authors also note "the limitations of LLMs in bug detection".

So keep the roles clear. In practice, AI proposes tests and inputs, while a software engineer checks that each test pins a real behaviour and adds the edge cases found earlier. The same person decides which captured outputs are bugs to fix later. Once the pinning tests pass, every later change is something you can verify rather than trust.

Should each part be retained, refactored, rebuilt or retired?#

Decide per component, not per system: retain what is stable, refactor what is sound but hard to change, rebuild what cannot carry the business forward, and retire what nobody uses. A rewrite is one of four answers, and usually the right one for only a few parts.

Retain, refactor, rebuild or retire: what each means and a typical legacy example. Source: strategy names from AWS Prescriptive Guidance, migration strategies, accessed 29 September 2026.

DecisionFor a part that isTypical legacy example
RetainStableA billing batch that passes its pinning tests
RefactorSound but hard to changeAn old Laravel or Java module on an unsupported framework
RebuildUnable to carry the business forwardAn AngularJS front end or a tangled monolith module, behind a facade
RetireUnusedReports, screens and endpoints the logs show nobody uses

The categories are not new. AWS's guidance on migration strategies, accessed on 29 September 2026, lists retire, retain, repurchase and refactor among its options. In particular, it defines retire as the choice for apps "that you want to decommission or archive". Also, it says you might retain an app "because it requires a detailed assessment and plan prior to migration".

Here is how the four answers look in a real estate of old systems:

  • Retain a stable billing batch that passes its pinning tests. Patch it and leave it.
  • Refactor an old Laravel or Java module with sound logic on an unsupported framework. Upgrade it in place.
  • Rebuild an AngularJS front end or a tangled monolith module. Build the new version behind a facade.
  • Retire reports, screens and endpoints the logs show nobody uses.

Because this choice needs the map from step one and the tests from step two, the assessment comes before the commitment. A team that signs up for a full rewrite first has answered the question for every part at once, usually without the evidence.

How does incremental modernization keep the business running?#

A facade routes each request to the old or the new code, so you move one slice at a time and can route back if a slice misbehaves. Microsoft's strangler fig pattern was updated on 2 June 2026. It puts it plainly: "the façade incrementally shifts requests from the legacy system to the new system."

DiagramA facade sends each request to the legacy system or the new one, and can route a slice back. Source: Microsoft Azure Architecture Center, strangler fig pattern, 2 June 2026.

The AWS version of the pattern, accessed on 29 September 2026, gives the reason. It warns that "a big bang migration approach is risky". With a facade, the old system keeps serving everything you have not moved yet. Meanwhile, the new one grows one route, one screen or one API at a time.

These slices are where AI speed is real, because each one is small, repeated and checked by tests you already wrote. Google reported this shape of work in a case study by Ziftci et al., published on 13 April 2025. Across 39 migrations, three developers over twelve months submitted 595 code changes with 93,574 edits.

Show data table
Bounded, repeated changes, each reviewed by a developer. The time saving is the developers' own estimate. Source: Ziftci et al., Google (arXiv 2504.09691), 13 April 2025.
Item Share (%)
Code changes generated by the LLM 74.45%
Edits generated by the LLM 69.46%
Estimated time saved (developer estimate) 50%

The model generated most of the submitted changes, and developers still reviewed and submitted every one.

Figure Bounded, repeated changes, each reviewed by a developer. The time saving is the developers' own estimate. Source: Ziftci et al., Google (arXiv 2504.09691), 13 April 2025. Ziftci et al., Google (arXiv 2504.09691), 13 April 2025

The model generated most of those changes, and the developers estimated they saved half their time. However, people still found, reviewed and submitted every change. In short, that is the model to copy: AI does the repeated edit, and a software engineer owns each slice from start to rollback.

Which support dates should set the order of the work?#

PHP 8.2 security support ends on 31 Dec 2026, about three months from now, so a Laravel application on it should start its assessment before its framework and runtime both stop receiving fixes. The date comes from PHP's supported versions page, accessed on 29 September 2026. Meanwhile, Laravel's release notes show that Laravel 11 security fixes ended on March 12th, 2026.

That is the worked example. Take a Laravel 11 app on PHP 8.2, assessed in September 2026. The framework is already out of security fixes, and the runtime has about three months left. In that window, the team maps the system, pins its behaviour and moves the first slice. Therefore, the first slice is the upgrade path, not the ugliest module.

Try it
Pick a stack

Dates from each publisher's own support page, counted from 29 September 2026.

The date your runtime or framework stops receiving security fixes.

31 Dec 2026
3

PHP 8.2: 3 months left#

Goes first

5 of the 8 listed stacks have a shorter runway. Parts on those go ahead of this one, however messy this code is.

Worked example: whole months of security support left from 29 September 2026
AngularJSJanuary 2022-56
Java SE 8 Premier SupportMarch 2022-54
Laravel 11March 12th, 2026-6
Node.js 20Mar 24, 2026-6
.NET 8 and .NET 9November 10, 20261
PHP 8.231 Dec 20263
Laravel 12February 24th, 20274
PHP 8.331 Dec 202715
Months of security support left from 29 September 2026. The shortest runway goes first. Pick a stack or enter your own support end date. Arithmetic from the published dates of The PHP Group, Laravel, OpenJS Foundation, Microsoft, the AngularJS project and Oracle, accessed 29 September 2026. Modelled, not measured.

What do the other legacy stacks look like?#

The same logic runs across the common stacks. The Node.js EOL page says an end-of-life version "will no longer receive updates, including security patches". Node.js 20 reached that status on Mar 24, 2026. Also, Microsoft's .NET support policy lists November 10, 2026 as the end of support for .NET 8 and .NET 9.

Older front ends are further along. The AngularJS project states: "AngularJS support has officially ended as of January 2022." Finally, Oracle's Java SE roadmap shows Premier Support for Java SE 8 ended in March 2022. Its Extended Support runs until December 2030.

In short, code quality still matters, but runway sets the order. A part on a runtime with months left goes ahead of a messier one on a supported stack. Put your own end date into the calculator, and the first slice usually picks itself.

What must software engineers keep control of?#

Architecture, business logic, validation, security and the production release stay with software engineers, because those are the decisions a wrong answer cannot be rolled back from cheaply. AI can draft all of them, but it should approve none of them.

What AI may draft and who approves it, for each of the five decisions. Source: method from this article; review model from Ziftci et al., Google (arXiv 2504.09691), 13 April 2025.

DecisionWhat it coversWho approves
ArchitectureThe target design, the facade and the data boundariesA software engineer signs it off
Business logicEvery rule recovered in step oneA software engineer, with the business owner
ValidationThe pinning tests and the pass rules for each sliceA person decides which outputs are bugs
SecurityLogins, secrets, dependencies and data handling in every changed fileA software engineer reviews it
Production releaseThe cutover and the rollback planA named person owns each release

First, the 2025 Google result shows why this split works. The model generated 74.45% of the code changes. Yet three developers found the change sites, reviewed every diff and submitted the work themselves.

Second, the 2025 METR trial shows the other side. There, trust without measurement ran the wrong way, since "AI tooling slowed developers down" while those developers believed the opposite.

In practice, write down who approves what before the first slice moves:

  • Architecture: the target design, the facade and the data boundaries. A software engineer signs it off.
  • Business logic: every rule recovered in step one. A software engineer confirms it with the business owner.
  • Validation: the pinning tests and the pass rules for each slice. A person decides which outputs are bugs.
  • Security: logins, secrets, dependencies and data handling in every changed file. A software engineer reviews it.
  • Production release: the cutover and the rollback plan. A named person owns each release.

With that list in place, AI can work fast inside a fence. It drafts, suggests and repeats, while the people who carry the risk make the calls.

When is AI-assisted incremental modernization the wrong fit?#

If a component is small, unused or about to be replaced by a product you buy, retire or replace it rather than spending AI or tests on preserving it. AWS's guidance calls the buy route repurchase, "also known as drop and shop". For plain functions like email or ticketing, it is often the cheaper path.

DiagramWhen AI-assisted incremental work is the wrong fit, and the better route for each case. Strategy names from AWS Prescriptive Guidance, migration strategies, accessed 29 September 2026.

There are two other cases where this approach is the wrong tool. First, when nobody on the team can check the output, AI only makes changes faster than they can be reviewed. In that case, bring in someone who knows the stack to write the pinning tests by hand, then add AI once there is a reviewer.

Second, when the change is a known, repeated upgrade, a rule-based tool often beats a model. For PHP, for example, fixed refactoring rules apply the same upgrade the same way every time. The guide to Rector for PHP migrations covers that route. Since the rules are fixed, the result repeats in a way model output does not.

In short, use AI where there is real reading or repeated editing to do, and a person to check it. Where there is neither, pick the simpler tool.

What should you do before committing to a rewrite?#

Commission an assessment of the existing codebase first: the rules, the dependencies, the tests you can write today and the support dates, then decide. That assessment is the cheapest step in the whole program, and every later choice depends on it.

A good assessment answers five questions. What does the system do, including the rules nobody wrote down? What does it depend on, and what depends on it? How much of its behaviour can you pin with tests today? Which parts should you retain, refactor, rebuild or retire? And how many months of support does each runtime have left? With those answers, you know where a rewrite is right and where it is not.

To modernize legacy code with AI safely, keep this order and keep people in charge of the decisions. For the next step, the guide on how to choose a legacy modernization partner covers what to ask whoever does the work. Then agentic development in the software lifecycle shows where AI fits when people keep review and release control. Finally, the enterprise software development guide covers the wider build, buy or modernize choice.

If you would rather have help, Atyantik offers AI-augmented development with software engineers in control. It also does enterprise software work for systems that need a refactor or a rebuild. Either way, the support pages, the papers and a careful read of your own code are enough to start the assessment yourself.

Questions this post answers

How do you modernize legacy code with AI without breaking business logic?
Use AI to read the system and pin its current behaviour with tests first, then change it in small reviewed slices, because blind AI translation was correct for at most 47.3% of samples.
Can AI convert a legacy codebase to a new language on its own?
No. A study of 1,700 code samples found 15 categories of translation bugs, and a converter cannot tell which quirk is a bug and which is a business rule. A model that converts line by line either copies the quirk or quietly drops it.
What should you do before committing to a rewrite of legacy code?
Commission an assessment of the existing codebase first: the rules, the dependencies, the tests you can write today and the support dates, then decide. That assessment is the cheapest step in the whole program, and every later choice depends on it.

Keep reading