Agile Was Built for a Different Software Constraint. What Comes Next?
In short
- Agile was designed when developer time was the scarcest resource. With agents handling execution, the bottleneck moves to decisions, context and verification.
- Sprints, story points and refinement are the practices under the most pressure. Sprint reviews and constant communication matter more.
- The Agile Manifesto's principles still hold. What's changing is the rhythm around them.
Agile, and Scrum in particular, grew up in a world where developer time was the scarcest thing on any software project. Most of the framework exists to protect that time: sprints, story points, backlogs and daily standups were all designed around it.
That world is changing. On AI-native projects, agents now handle a large part of the execution, and the bottleneck has moved. Today it sits in deciding what to build, checking that it's right, and getting enough business context to make those calls.
I don't think this means Agile is dead, but the future of Agile looks different from its past. The 12 principles of the Agile Manifesto hold up better than ever, because we iterate and change direction faster than we used to. Scrum's machinery is another story. Its cadences, ceremonies and units of measure all assumed that human execution capacity was the limit.
What problem was Agile originally solving?
When the Agile Manifesto was written in 2001, building software was slow and expensive, and most of that cost was developer hours. So the process was designed to keep developers moving on the product. Short iterations kept work focused, estimates helped plan capacity, and standups cleared blockers before they cost the team a day.
Where is the bottleneck in AI-native software development?
When a feature that used to take a developer a week comes back from an agent in a few hours, execution stops being the scarce resource.
In our projects, the hard part now happens before and after the code. Before, it's deciding what to build and making sure the team, people and agents, has enough business context to build the right thing. After, it's knowing whether what got built actually solves the problem. A big part of the AI development process now happens before anyone opens an editor.
Agents are only as good as the context they get, and they need clear decisions to point them in the right direction.

When agents take most of the build, the constraint moves to the stages before and after it. Diagram comparing where the software delivery bottleneck sits: in classic Agile, building was the narrow stage; in AI-native delivery, deciding and verifying are.
Does a two-week sprint still make sense?
Before, a developer would usually take two or three features in a sprint. With agents, some of those features are ready in a few hours. Some need a few rounds of iteration before they land, but the time scale is different.
So the sprint backlog has started to look more like a Kanban board, a flow of work that agents and developers pick up as soon as it's ready.
Clients still want to see progress, and demos still make sense. But there's no reason to sit on a working feature for two weeks because that's when the review is scheduled. On some projects, the sprint is getting closer to one day.
What replaces story points?
At Rootstrap, we've replaced story points with capability-based estimation. Story points and velocity measure human effort at a granular level, and that stops making sense when agents can build a 13-point feature and a 100-point feature in about the same time.
A capability is a business function, like User Management, Ecommerce Checkout or Document Ingestion. It's something the system does, a unit you could scope and deliver on its own, made of stories with acceptance criteria and evals. We think it's the right level of detail for AI-native delivery because it's a unit of business value, so clients can recognize the result an estimate describes. Every capability carries a class, a tier and a size, which gives us a consistent way to compare work across projects.
The market hasn't agreed on a standard yet, but that consistency is what makes the next step possible. Every capability delivered through our new Delivery System generates telemetry. Over time, that data shows how different types of work actually perform in real projects. We want to use it to keep improving how we forecast delivery and, eventually, move toward probabilistic proposals: a range of likely scenarios grounded in real delivery data instead of assumptions.
How does backlog prioritization change?
We still use a backlog, and prioritization still matters, but at a different level of detail. When small items take hours, ranking every story against every other one doesn't add much. Prioritizing at the epic level is usually enough.
Scope discipline matters as much as it ever did. An MVP still needs a limited set of features. Agents can build more, and build it faster, but people get overwhelmed when a product grows in too many directions at once.
What happens to handoffs?
Handoffs used to come with waiting time built in. One role finished, the next one picked up, and some context got lost along the way.
Two things changed. Documentation became more important, because it's what agents build from. It also became much easier to produce and use. We used to need separate meetings with each role on a project to share context. Now we can put the documents and conversations into an LLM and get a summary tailored to each person.
Onboarding works the same way. A developer joining a project still has to understand the business, and there's no shortcut for that. The technical setup takes much less time, though, since agents do most of the coding.
When does verification happen in AI software delivery?
If execution isn't the bottleneck, verification is, and it doesn't wait for the end of the sprint anymore.
On our projects it happens all the time and at several levels. Agents check their own work, developers review it, QA agents test it, and QA testers validate it. Features are built continuously, so they get checked continuously.
Decisions need attention too. A fixed weekly checkpoint for open decisions can work, as long as it doesn't hold up delivery. My view is that decisions will simply need to be made faster. When everything else speeds up, a decision that waits a week is what slows the project down.
Which Agile ceremonies matter more, and which matter less?
Refinement is the one I see losing weight first. If documentation is good and conversations are captured, most of what refinement used to do is already covered. Long written feature descriptions matter less too. A Product Manager can show what a feature could look like with a prototype faster than they can explain it in a ticket, and AI turns that prototype into instructions for the agents.
Planning gets harder to use as a forecast. When the output of a cycle varies this much, predicting how many features will land in a given window is tricky.
Standups stay. People still need to align on who does what. Blockers come up faster with agents, like an agent loop that stalls or a decision nobody has made yet, but they also get solved faster. That means the team needs to talk more often.
Sprint review becomes more important. Teams move faster, so clients want to see progress more often, and reviews happen more frequently.
Retrospectives are shifting too. When work ships in releases, a post-mortem after each release can teach the team more than a retro every two weeks.
Eliana Cibulis, a Project Manager at Rootstrap, saw all of this on one of her projects:
"Once we started using multi-agent workflows, our refinement sessions became less relevant. The product manager began creating detailed specs in Notion using AI and would share them directly with the developer taking on the capability. They would then review the requirements together to make sure everything was clear, and, when necessary, schedule a 1:1 to discuss any questions in more detail.
Our retrospectives also started to lose some of their value as the team adapted to this new way of working. Instead, we agreed to hold a post-mortem after the latest release to reflect on what worked well, what could be improved, and what we could learn from the process.
Dailies and demos, however, remained an important part of our workflow. We kept the daily meetings because they were still valuable for understanding what everyone was working on and identifying any blockers that needed attention.
We also kept demos because we were working on releases, each with a specific scope. The demos gave us an opportunity to show the client our progress and gather feedback before moving forward with the release."
The goal of each ceremony is still valid. Dailies keep the team in sync and keep the team spirit going. Retrospectives and post-mortems are where we learn and improve. Reviews show progress. Refinement defines the product and sets priorities. The difference is how often each one happens and how long it takes.

Refinement and planning lose weight, standups hold, and sprint reviews happen more often. Slope chart of Agile ceremonies from classic Scrum to AI-native delivery: refinement and planning matter less, standups and retrospectives hold, sprint review matters more.
What doesn't change
Building products people actually want still takes iteration. You get there by approximation: build something, show it, learn, adjust. That was Agile's central idea, and it still is. AI makes each loop faster, so teams get more chances to get it right.
Others are seeing the same shift
Plenty of people are writing about Agile and AI right now. Sam Gao, a CMMI lead appraiser who runs his own software company, reaches a very similar conclusion in his proposal for what he calls AI-Scrum. He argues that Scrum's principles hold, but its ceremonies, team structures and time horizons need a redesign. He's also seeing sprints shrink toward daily cycles, prototypes take the place of written requirements, and story points lose meaning when a large task and a small one finish in similar time.
On estimation, we've taken a different route. Gao has moved toward counting completed tasks as a proxy for progress, and we're testing capability tiers. Both models are still early. They start from the same place, though: once agents do most of the execution, a metric based on human effort stops reflecting what's actually happening.
What comes next
There's no new framework yet. What replaces Scrum might be a new version of it, or something different altogether. Some people already call this post-Agile development. For me, the future of Agile is its principles running on a new operating rhythm, because the original setup was built for a constraint that's going away.
If you're choosing a delivery partner through this shift, these are the questions worth asking.
FAQ
What is the future of Agile in AI-native software development?
Agile's principles stay, and its practices get faster and lighter. With AI agents handling execution, sprints shrink toward continuous flow, story points give way to new estimation units, and verification happens all the time instead of at the end of a sprint.
Is Agile dead in the age of AI?
No. The 12 principles of the Agile Manifesto still apply, and faster iteration makes them even more relevant. The practices under pressure are some of Scrum's: sprint length, estimation, and how often ceremonies happen.
What replaces story points in AI-native software development?
There's no industry standard yet. At Rootstrap, we're testing capability-based estimation. Each capability gets a class, one of five complexity tiers (C1 to C5) and one of two sizes per tier, measured in weeks of a full-stack developer working with agents.
How long should a sprint be when using AI agents?
It depends on the project, but the two-week default is losing ground. With agents delivering some features in hours, some teams are moving toward a continuous flow of work, with demos as often as daily.