Two review queues for the same reviewer. Before agents, a short stack of pull requests takes 20 hours of review a week. With agents, pull requests rise 98% and each review takes 91% longer, so the week grows to 75.6 hours, a multiplier of 3.78.

The agentic SDLC: how AI agents change specification, verification and governance

AI agents can now take a task, work on a branch and hand back a change. The hard part moved. It now sits in what you write before an agent starts and in how you prove its work before anyone merges it.

What is the agentic SDLC, and what does it change?#

The agentic SDLC hands execution at each stage to AI agents and keeps people on specification, verification and sign-off, because pull requests rose 98% while review time rose 91%. Those two figures come from a May 2026 paper on arXiv that pooled telemetry across more than 10,000 developers. In the same data, delivery metrics stayed flat.

Show data table
More pull requests, longer reviews, flat delivery. Source: arXiv 2605.01160, 1 May 2026, telemetry across 10,000+ developers.
Item Value
Pull requests opened 98%
Review time 91%

Pull requests nearly doubled, and each review took almost twice as long.

Review telemetry More pull requests, longer reviews, flat delivery. Source: arXiv 2605.01160, 1 May 2026, telemetry across 10,000+ developers. arXiv, 1 May 2026

What separates it from AI-assisted coding is the unit of work. For example, an assistant suggests a line or a function while you type. By contrast, an agent takes a whole task, reads the repository, runs tools and returns a change. An April 2026 survey of agentic systems puts the shift in one phrase: "from code generation to delegated execution under human supervision".

Because of that, delegation changes the joints of the lifecycle more than the stages. Each stage now has a handoff with four parts. It names the input an agent is handed and the artifact it emits. Then it names the check that proves the artifact without trusting its author, and the person who signs. Every section below fills in those four parts for one stage of the agentic SDLC.

The four-part handoffThe handoff at every stage of an agentic SDLC: what the agent is handed, what it emits, the check that proves it, and who signs. Source: Atyantik Technologies, from the stage contract in this post.

How many review hours does an agent's extra output cost?#

A team reviewing for 20 hours a week needs about 75.6 hours once pull requests rise 98% and each review takes 91% longer. The arithmetic is short, because volume and review time multiply. Using the rises in the May 2026 arXiv telemetry, the sum is 20 × 1.98 × 1.91, which comes to 75.6 hours.

Your review load with agents

Put in your team's weekly review hours; the two rises start at the May 2026 arXiv telemetry, and you can change either one.

Your review load before agents

measured on this projectan assumption, change it

Only review hours are counted. Rework, waiting time and the checks that could shorten a review are left out, because they differ for every team.

Review hours per week after agents

75.6

Times the review load before agents, to one decimal
3.8

Arithmetic from arXiv 2605.01160, 1 May 2026. Modelled, not measured.

Show data table
Baseline 1.0, pull request volume only 1.98, review time only 1.91, both 3.78. Arithmetic from arXiv 2605.01160, 1 May 2026.
Item Value
Baseline 1 x
Pull request volume only 1.98 x
Review time only 1.91 x
Both 3.78 x

Each rise alone roughly doubles the review load, and together they nearly quadruple it.

Why the rises multiply Baseline 1.0, pull request volume only 1.98, review time only 1.91, both 3.78. Arithmetic from arXiv 2605.01160, 1 May 2026. Arithmetic from arXiv 2605.01160, 1 May 2026. Modelled, not measured

In short, the May 2026 telemetry implies a review load about 3.78 times larger unless something absorbs it. As a result, that multiplier is the number to take to a planning meeting. However, few teams compute it, since agents are bought on how fast they write.

Two levers bring that load down. First, volume: an agent handed a vague task opens more pull requests to get one right, so a sharper task means fewer tries. Second, time per review drops once automated checks have proved a change meets its written criteria. Then the reviewer reads for design and risk, not for whether it works.

Therefore, every stage that follows is a way to win back review hours. Specification cuts the volume, and then independent checks cut the time per review. Finally, governance makes sure someone still owns the result.

Why did the speed forecasts miss the measured result?#

In a randomised trial on mature projects, developers forecast AI would cut task time by 24%, yet tasks took 19% longer. The trial was published on arXiv in July 2025 by METR. It followed 16 experienced developers through 246 tasks on projects they had worked on for about five years.

Show data table
Forecast against measured task time, 16 developers, 246 tasks. Source: arXiv 2507.09089 (METR), July 2025.
Item Value
Economics experts forecast: shorter 39%
ML experts forecast: shorter 38%
Developers forecast: shorter 24%
Developers' estimate afterwards: shorter 20%
Measured: longer 19%

Every group expected shorter tasks, yet the measured tasks took 19% longer.

Forecast against measured Forecast against measured task time, 16 developers, 246 tasks. Source: arXiv 2507.09089 (METR), July 2025. arXiv (METR), July 2025

The gap between belief and result is the striking part. In the July 2025 paper, experts in economics forecast 39% shorter tasks and experts in machine learning forecast 38% shorter. Also, even after the work, the developers believed AI had made them 20% faster. However, the clock said 19% slower.

That result does not mean agents fail everywhere. For instance, the May 2026 review cites controlled studies with 20% to 56% gains on well-scoped tasks. Instead, it says the setting decides, and in particular task scope, codebase maturity and developer experience all move the result.

A September 2026 synthesis on arXiv names where the gains go. It found that "those gains attenuate sharply between writing code and shipping reliable software". In practice, review, testing, security and deployment remain the slow stages. As a result, the controls around an agent matter more than its raw speed, and each stage needs a written contract.

What must a specification hold before an agent reads it?#

An agent reads a specification literally, so acceptance criteria, boundaries and the checks that prove done must be written before the task is dispatched. For instance, a person fills gaps in a ticket with judgment and a quick question. By contrast, an agent fills them with a guess, and the guess comes back as a pull request someone has to review.

The May 2026 paper on the productivity paradox puts it flatly: "Specification discipline, not model capability, is the binding constraint" on dependable AI-assisted software. Also, a spec-driven development paper from the same month makes the case from the other side. It argues agentic speed "relocates discipline upstream into specification precision", along with explicit gates and auditable provenance.

In practice, a task an agent can take has four fields:

  1. Outcome: What changes for the user or the system, in one or two sentences.
  2. Acceptance criteria: Each one is testable, so a check can pass or fail it without a person reading code.
  3. Boundaries: Which files, services and data an agent may touch, and which it must leave alone.
  4. Proving checks: The tests or commands that decide done, written before the work starts.

However, the fourth field is the one teams skip. Without it, an agent writes its own tests and passes them. Then the reviewer has to judge the code and the tests together, and that is where the review hours go.

Where the proving checks come fromWhere the proving checks come from decides what the reviewer reads: the written task's checks, or the agent's own tests beside its code. Source: Atyantik Technologies, from the four-field task in this post.

So the handoff for this stage is clear: the input is the business need, and the output is the four-field task. Before dispatch, a second person reads the task for gaps, and then the product owner signs it.

How does implementation change when an agent works on a branch?#

Benchmark resolve rates rose from 1.96% to 78.4%, so implementation is now a bounded agent session on a branch whose output is a candidate change. In particular, the April 2026 survey on arXiv reports that rise on SWE-bench Verified, a set of real GitHub issues, between October 2023 and April 2026.

Best reported resolve rate on SWE-bench Verified1.96% to 78.4%

1.96%

October 2023

78.4%

April 2026

Agents went from solving almost no real GitHub issues to solving most of them.

Best reported resolve rate on SWE-bench Verified (percent of SWE-bench Verified tasks resolved)
Optionpercent of SWE-bench Verified tasks resolved
October 20231.96%
April 202678.4%

Source: arXiv 2604.26275, 29 April 2026

Meanwhile, the tools now work in sessions with a start and an end. For example, GitHub's documentation for Copilot cloud agent says it can "make code changes on a branch" inside "its own ephemeral development environment, powered by GitHub Actions". GitHub also sets "a maximum execution time of 59 minutes" per session. In addition, Copilot "can open exactly one pull request to address each task it is assigned".

Terminal agents follow the same session pattern. Anthropic's overview of its terminal coding agent says you can "Edit files, run commands, and manage your entire project from the command line". Also, it works with git: it "stages changes, writes commit messages, creates branches, and opens pull requests".

So the shape of implementation changes: a session takes the four-field task, runs inside a time box and ends with a branch. Instead of a merge, that branch is only a candidate. Although capability rose fast, trust did not rise with it, so the session's output waits for the next stage.

Here is the handoff for implementation. First, the input is the task, and the output is a branch with one pull request. The check is the time box plus branch protection, and the person who dispatched the task owns the session.

How do you verify code when the author is an agent?#

Only 33% of developers trust the accuracy of AI output and 46% actively distrust it, so verification must not depend on the coding agent's own tests. Those figures come from the Stack Overflow 2025 Developer Survey, which also found only 3% "highly trust" the output.

Show data table
Developers' trust in the accuracy of AI tools. Source: Stack Overflow 2025 Developer Survey.
Item Value
Actively distrust the accuracy of AI tools 46%
Trust the accuracy 33%
Highly trust the output 3%

Nearly half of developers actively distrust the accuracy of AI output, and only 3% highly trust it.

Trust in AI accuracy Developers' trust in the accuracy of AI tools. Source: Stack Overflow 2025 Developer Survey. Stack Overflow, 2025 Developer Survey

Therefore, the fix is to split writing from release. In particular, the spec-driven development paper from May 2026 argues for "separation between synthesis and release authority". In short, an agent may write the change, but it may not decide the change is ready.

Three rules make that split real:

  1. The checks come from the specification: They were written before the session started, so no agent shaped them to fit its code.
  2. The pipeline runs them: They run on the branch in a job no agent can edit, protected by the same rules as any other release gate.
  3. An agent's own tests are evidence: They can show intent, but a green run on tests an agent wrote proves only that it agrees with itself.

Also, review changes its job. Once the checks pass, the reviewer reads for design, risk and fit with the rest of the system. Since that is the work a person does better than a script, the review hours from the calculator get spent where they count.

So the handoff for verification runs like this. The input is the branch, and the output is a pass or fail on the written checks. The protected pipeline is the check, and a named reviewer signs the merge.

Who signs what at each stage of an agentic SDLC?#

Governance in an agentic SDLC is a named human signature per stage plus a provenance record of what the coding agent was given and produced. Accountability never moves to an agent, because an agent cannot answer for a failure in production.

Existing frameworks already cover most of this governance work. For instance, NIST published its Secure Software Development Framework in February 2022. It describes "a core set of high-level secure software development practices that can be integrated into each SDLC implementation". Since an agentic SDLC is one more SDLC implementation, the same practices apply to what agents write.

Provenance is the record that ties a change to how it was made. The SLSA requirements say the build platform is "responsible for generating provenance describing how the package was produced". Also, they say the producer "MUST distribute provenance to artifact consumers". For an agent-made change, the record should also name the task it was handed and the session that produced it.

For the wider picture, NIST released its AI Risk Management Framework on January 26, 2023, and it is "intended for voluntary use". In practice, it gives leaders a shared way to talk about AI risk across the whole company, above the level of one pipeline.

Here is the full contract, one row per stage:

The agentic SDLC stage contract: what the agent is handed, what it emits, the check that does not trust the author, and who signs. Source: Atyantik Technologies, from the stage sections of this post; provenance duty per SLSA v1.0 requirements.

StageAgent is handedAgent emitsCheck that does not trust the authorWho signs
SpecificationThe business needA draft four-field taskA second person reads it for gapsProduct owner
ImplementationThe signed taskA branch and one pull requestTime box and branch protectionPerson who dispatched it
VerificationThe branchResults of its own testsPipeline runs the checks from the taskNamed reviewer
ReleaseThe merged changeRelease notes and a changelog draftBuild provenance from the platformRelease owner

Which pages do you build the first agent lane from?#

Three documentation pages cover the first lane: the cloud agent's session and branch model, the terminal agent's commands, and the provenance requirements. Read them in this order, and take one thing from each:

  1. About Copilot cloud agent, from GitHub. Take the session model: one branch, one pull request per task, a 59-minute hard limit, and the timeout-minutes setting in copilot-setup-steps.yml for a shorter box.
  2. Anthropic's terminal coding agent overview, from Anthropic. Take the command surface: starting a task from the terminal, and running in CI with GitHub Actions or GitLab CI/CD.
  3. SLSA requirements, version 1.0, from the OpenSSF SLSA project. Take the provenance duties: the build platform generates it, and the producer distributes it.

Either kind of agent works for this lane. In short, what matters is that the session, the branch and the record are defined before the first task goes out.

What is the smallest agent lane a team can run this week?#

The smallest lane is one time-boxed agent session per written task, a branch as its only output, and a pipeline check the coding agent cannot edit. It takes three steps: first cap the session, then hand over the task with its specification, and finally verify on the pipeline with checks no agent owns.

bash
# 1. Cap each Copilot cloud agent session below GitHub's 59-minute hard limit.
mkdir -p .github/workflows
cat > .github/workflows/copilot-setup-steps.yml <<'EOF'
on: workflow_dispatch
jobs:
  copilot-setup-steps:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v4
EOF

# 2. File the written task as an issue, then select Copilot as its assignee on GitHub.
gh issue create --title "Task 142: export invoices as CSV" --body-file docs/tasks/142.md

# 3. In the pipeline, run the checks written into the task, from a path agents may not edit.
bash checks/142.sh

Step 3 is the one that matters most in the lane. Therefore, protect the checks/ path with code owners, so a change to it needs a person's approval. Then an agent can propose a new check, but it cannot weaken an old one to make its own work pass.

When is an agentic SDLC the wrong fit for a stage?#

Delegating a stage is the wrong fit when no check can prove the output, or when experts on a mature codebase already work faster than review allows. In those cases, keep the stage with people and give agents smaller jobs inside it.

No check can prove it. Some work has no test that decides done: a tax rule with legal weight, an architecture choice, a migration with no way back. Instead, hand an agent the first draft of the design note or the test plan. Then a person makes the call, with the draft in hand.

Experts on a mature codebase. The July 2025 trial on arXiv measured a 19% slowdown for developers who knew their projects well. For that work, use agents for the parts with clear checks, such as tests for untested code or updated documentation. Meanwhile, the core change stays with the expert.

Almost-right output that costs more to fix. In the Stack Overflow 2025 survey, 66% of developers named one frustration above all others. It was "AI solutions that are almost right, but not quite". When a stage keeps producing that kind of output, the task is too loose. In that case, go back to the specification stage and tighten it before you dispatch again.

When to delegate a stageWhen to delegate a stage, and what to hand an agent inside a stage that stays with people. Source: Atyantik Technologies; slowdown per arXiv 2507.09089 (METR), July 2025; almost-right frustration per Stack Overflow 2025 Developer Survey.

Where should a team go from here?#

Pick one stage, write its handoff contract, size its review load, and delegate it only once the check that proves it exists. In practice, small, well-tested changes are the usual first choice, because their checks already exist.

For the people side of the shift, read how AI changes developer teams. For where the lane sits in your process, see software development methodologies. Verification has its own depth in AI and ML in software testing. Also, modernising legacy code with AI shows agents at work on the kind of mature codebase the trial studied.

If you want the lane built with you, there are two routes. Our AI-augmented development team delivers agent-written code under specification and review controls. Meanwhile, our DevOps and CI/CD team builds the pipeline checks that absorb the extra pull requests. Still, the three documentation pages above are enough to build the first lane of an agentic SDLC on your own.

Questions this post answers

What is the agentic SDLC?
The agentic SDLC hands execution at each stage to AI agents and keeps people on specification, verification and sign-off, because pull requests rose 98% while review time rose 91%. Those two figures come from a May 2026 paper on arXiv that pooled telemetry across more than 10,000 developers.
How many review hours does an agentic SDLC add?
A team reviewing for 20 hours a week needs about 75.6 hours once pull requests rise 98% and each review takes 91% longer. In short, the May 2026 telemetry implies a review load about 3.78 times larger unless something absorbs it.
When is an agentic SDLC the wrong fit for a stage?
Delegating a stage is the wrong fit when no check can prove the output, or when experts on a mature codebase already work faster than review allows. In those cases, keep the stage with people and give agents smaller jobs inside it.

Keep reading