Does Adopting AI Really Make Teams Faster?
Developers who feel faster, teams that stay the same

Lately, many organizations have been adopting AI almost as if it were a race. At Cloud Next in April 2026, Google announced that 75% of all new code at Google is now generated by AI and approved by engineers. Just a year and a half earlier, in its October 2024 earnings call, that number was only a little over a quarter.
With numbers like these showing up in the news every quarter, other organizations can hardly sit still, and it has become a familiar sight to see internal AI coding tool adoption rates or weekly active user ratios written into quarterly goals.
Yet among these organizations, it seems not that many actually measure whether their teams got faster because of AI. In fact, in a report published in 2025 by LeadDev, a community for engineering leaders, 82% of organizations said they weren’t measuring the impact of their AI coding tools.
At Toss, where I work, we’ve built metrics for productivity and application stability and are trying to observe whether those metrics move in a good direction as we adopt AI. The reason is simple: you can’t just adopt AI blindly without even measuring whether things got better or actually regressed.
In fact, I touched on this problem not long ago in a post called On Properly Defining Ambiguous Problems. The point was that if it takes 3 weeks for a feature request to reach production and only 4 of those days are spent writing code, then even if AI cuts the writing time in half, 21 days only shrink to about 19. Push the same calculation all the way, and even if AI brings code-writing time down to zero, a 21-day lead time only drops to 17 days, a measly 19% or so.
Many developers say that since adopting AI, their productivity has gone up or their team works faster. But is that really true?
In this post, I’ll first look at studies on how much faster developers and teams using AI have actually become, and then translate the old saying that a chain is only as strong as its weakest link into the language of the theory of constraints and Little’s law to work out why those results come out the way they do.
After that, I’ll look at why organizations don’t measure this, and then lay out what needs to change. Finally, we’ll go over three metrics worth measuring first for a team that has adopted AI.
Feeling Faster and Being Faster
Most developers who have used AI feel that they work faster. In fact, in DORA’s 2025 survey, more than 80% of respondents said AI had increased their productivity.
But feeling faster and actually being faster are two different things. So in this section, I want to look at the two separately. First, we’ll look at an experiment that compared how fast developers felt they were with how fast they actually were, and then use large-scale surveys to see what happened to the team as a whole when individuals got faster.
Developers Believed They Had Gotten Faster
In July 2025, the AI research organization METR published a rather surprising experimental result. They gave 16 experienced developers, who had contributed to large open-source projects for over 5 years on average, 246 real issues, randomly assigned each issue to a condition where AI was allowed or not allowed, and measured how long each took.
Before the experiment, the developers expected AI to cut their working time by 24%, and even after it ended, they felt it had cut their time by about 20%. But when the time was actually measured, it turned out that the tasks done with AI took 19% longer. Economists and machine learning experts had predicted speedups of about 39% and 38% respectively, so they were wrong about the direction itself.
Of course, it’s hard to conclude from this one experiment that AI makes developers slower. With only 16 participants, the sample is very small, the setting was the special case of large repositories the developers had known for years, and the models used were from early 2025.
In fact, in a follow-up experiment that METR expanded to 57 participants and published in February 2026, the results pointed toward AI users being faster, but the margin of error was wide enough that a slowdown couldn’t be ruled out, and METR itself called it very weak evidence. That’s because quite a few participants simply left out tasks they didn’t want to do without AI.
What caught my attention in this experiment, though, wasn’t whether developers got faster or slower but that there was a sizable gap between how they felt and what actually happened. If developers felt faster even when they were slower, then a survey asking them how much AI improved their productivity isn’t as reliable a source of data as one might think. And yet a survey like this is exactly how many organizations check the impact of AI adoption, which is the unfortunate part.
So why does this illusion happen? I think the reason can be found in how the pattern of spending time changes when working with AI.
When you write code on your own, you spend a long time racking your brain, and all that time is remembered as hard work. When you work with AI, on the other hand, code fills the screen in seconds, and the rest of the time goes into reading that code, fixing it, and prompting again. The time spent thinking went down, but the time spent reviewing and fixing went up instead.
The problem is that this time spent reviewing and fixing isn’t short, either. In Stack Overflow’s 2025 developer survey, 66% of respondents named AI answers that are almost right but not quite as their biggest frustration, and 45% said debugging AI-generated code takes more time.
On top of that, AI-written code is often “plausible,” and finding what’s wrong in plausible code means reading all of it in the end, so you’re likely to burn more time on it than on code that’s simply wrong.
The interesting part, though, is that people don’t remember this time as time spent working. An hour of agonizing in front of a blank screen feels long, but an hour spent tweaking AI-generated code bit by bit can feel short because of the sense that something kept moving forward. This is my own interpretation and not a conclusion of the METR study, but I think it’s a fairly plausible explanation for why perception and reality can go in opposite directions.
Individuals Got Faster, but the Team Stayed the Same
Even if we grant that individual developers using AI became more productive, that doesn’t necessarily carry over to the team. In July 2025, Faros AI, a developer productivity analytics company, published a report analyzing real work data from 1,255 teams and more than 10,000 developers.
Developers on teams that used AI heavily clearly did more work. Completed tasks went up 21% and merged PRs went up 98%, so looking only at individual output, it nearly doubled.
But on the same teams, the time spent on PR review went up 91%, the average PR size grew 154%, and bugs per developer rose 9%. On top of that, the report concluded that at the company level, there was no meaningful correlation between AI adoption and improved performance.
A report published in 2024 by Google Cloud’s DORA research team showed a similar conclusion. For every 25% increase in AI adoption, documentation quality improved 7.5%, code quality 3.4%, and code review speed 3.1%, but the throughput at which teams actually deployed changes fell 1.5%, and delivery stability dropped 7.2%. Nearly every individual metric improved, yet the results of the team actually shipping product to users got worse.
Of course, the picture changed somewhat over a year. In DORA’s 2025 report, the relationship between AI adoption and delivery throughput turned positive, but stability was still negative. And in the end, the report summarized its findings like this.
AI doesn’t fix a team; it amplifies what’s already there. Strong teams use AI to become even better and more efficient. Struggling teams will find that AI only highlights and intensifies their existing problems.
In other words, AI doesn’t fix a team; it amplifies what’s already there. I think this sentence is a pretty accurate summary of the question this post is asking. AI can make individuals faster, but whether that turns into team speed depends on how the team was working to begin with.
So why the same AI makes some teams better and only magnifies problems for others seems to come down to what capabilities the team already had.
In fact, in its 2025 report, DORA laid out 7 organizational capabilities that amplify the effects of AI, including working in small batches, strong version control, and well-built internal platforms. And these were already considered the marks of a good team before AI came along.
I don’t think this is a coincidence. A team that works in small batches, reviews quickly, and has automated testing and deployment is a team whose stages outside of code writing are already running smoothly.
On such a team, code writing was likely the slowest stage, so when AI speeds it up, the whole team is likely to get faster. On the other hand, on a team where reviews pile up and deployments are unstable, code writing wasn’t the slowest stage, so speeding it up with AI only sends more work to the place that was already clogged. We’ll look a little more closely at why this difference arises in the next section.
For reference, in What Leaders Should Really Care About Isn’t Productivity, I wrote from the premise that the better you use AI, the more explosively individual productivity goes up, and the productivity I meant there was closer to the amount of code an individual produces per unit of time.
This post is more of a closer look at that premise. As the reports above showed, an increase in the amount of code individuals produce and an increase in what the team ships to users are less correlated than you might think.
A Chain Is Only as Strong as Its Weakest Link
The phenomenon of individuals getting faster while the team doesn’t isn’t actually a problem AI created for the first time; manufacturing worked it out decades ago. The old saying that a chain is only as strong as its weakest link is the core of the answer.
In this section, I’ll first look at the theory of constraints, which translated that saying into the language of the factory, and then point out that the claim “the team is fast” actually mixes 2 different meanings. Once this distinction is made, it becomes fairly natural to explain why the studies we saw came out the way they did.
Herbie Sets the Pace of the Line
In 1984, the Israeli physicist Eliyahu Goldratt published ”The Goal,” a business book written as a novel, and it contains one famous scene.
It’s the scene where the protagonist goes along on his son’s Boy Scout hike, watches the line keep stretching out and slowing down, and finally realizes that the whole line can only walk at the pace of Herbie, the slowest kid.
No matter how fast the kids in front walk, the time the line finally reaches its destination is the time Herbie, the slowest kid in the line, gets there. In this situation, the further ahead the faster kids get, the more the gap between them and Herbie just widens.
So the protagonist puts Herbie at the front of the line and has the other kids split up the things in Herbie’s backpack, raising the speed of the whole line. He didn’t make the fast kids faster; he made the slowest kid less slow.
Goldratt took this scene to the factory and named it the “theory of constraints.” The output of a factory made of a series of processes is determined by the slowest process, the constraint, and no matter how much you improve processes that aren’t the constraint, the factory’s overall output doesn’t go up.
The book also has a line that condenses this principle: an hour lost at the constraint is an hour lost for the entire system, and an hour saved at a non-constraint is a mirage. Even if a non-constraint process gets an hour faster, if it sits before the constraint it just finishes an hour early and waits an extra hour in front of the constraint, and if it sits after the constraint it just waits an extra hour for the constraint to hand over work, so either way the factory’s daily output stays the same.
In fact, the first problem the protagonist runs into in the book is measurement. His factory was focused on raising the utilization of each process, and since every machine was running nonstop, it looked like a very efficient factory on paper.
But the factory was actually piled high with inventory, and deliveries kept slipping. The parts the non-constraint machines were making nonstop were stuck lined up at the process that was the constraint. Every number measuring the efficiency of each process looked good, yet the amount the factory shipped to customers stayed the same, and this scene looks quite a lot like the Faros AI developer productivity report we saw earlier.
In the end, a software team is structured much like a factory. For a feature to reach users, it has to go through stages in order, such as planning to decide what to build, development to write the code, review where colleagues read the code, QA, and deployment.
And what AI is mainly speeding up right now is development and design. If this team’s Herbie wasn’t code writing or design work, then AI has merely made the kids at the front of the line walk faster.
Throughput Is Set by the Slowest Stage, Lead Time by Every Stage
When we say a team got faster, we’re often actually mixing 2 different things.
One is the team’s throughput. It refers to the number of features or changes the team ships to users in a week, and it’s like standing at the very end of the team’s process and counting how many features get deployed in a week.
Each stage has its own amount it can process, its capacity, but because the team’s throughput is measured at the very end of the process, its upper limit is set by the capacity of the slowest stage, just as in the Herbie story. For example, if planning prepares 5 features a week and development builds 10 but review can only handle 3, the team’s throughput ends up being 3 features a week.
The other is lead time. It refers to how long it takes from when a single feature request comes in until it reaches users, and it’s like following one piece of work from the start and measuring how many days it takes from going in to coming out. Lead time is the sum of the time spent in each process and the time spent waiting between processes, so it’s determined not by the single slowest process but by all of them together.
Since users look at how many weeks it takes for a bug they reported to get fixed and when the feature they requested will come out, they mostly look at lead time rather than throughput.
On the other hand, what an organization looks at when tallying how much it accomplished each quarter is closer to throughput. And the key point is that the individual output shown on AI adoption dashboards is neither of the two. (The odd part is that they show how many tokens you used, not how much work got done.)
So when talking about the impact of AI adoption, you first have to decide what got faster. An individual making more PRs only means one link in the chain got faster, and whether that turned into throughput and lead time is something you have to check separately.
That doesn’t mean speeding up a stage that isn’t the slowest one, the bottleneck, is entirely meaningless. Suppose a team whose bottleneck is review cuts QA, which comes after review, from 3 days to 1. Since the amount passing review stays the same, the throughput the team ships each week doesn’t go up, but the lead time for a feature to reach users drops by 2 days. Stages behind the bottleneck are always waiting for the bottleneck to hand over work, so there’s no queue in front of them, and the reduced processing time comes straight off the lead time.
The problem is that the same improvement can move lead time in opposite directions depending on whether it happens before or after the bottleneck. When a stage in front of the bottleneck gets faster, the extra work lines up in front of the bottleneck, and lead time can actually grow by that much waiting time.
And the key point is that on a team whose bottleneck is review, the code writing and design work that AI is speeding up happen to be exactly the stages in front of the bottleneck. We’ll work through this with numbers in the next section, but here, let’s first calculate the most optimistic case, assuming the throughput of the remaining stages stays the same.
The calculation I did in the September post was about exactly this lead time side. If code writing takes 4 out of 3 weeks, or 21 days, writing accounts for about 19% of the total. How much the whole speeds up when only part of it gets faster can be calculated with Amdahl’s law, formulated by computer scientist Gene Amdahl in 1967.
In the formula above, is the fraction of the whole taken up by the part that gets faster, is how many times faster that part gets, and is how many times faster the whole gets. It was originally made to figure out how much a whole program speeds up when part of it is parallelized, but it applies just as well to a team’s lead time under the assumption that the remaining stages stay the same.
Plugging in and gives an of about 1.1, which means 3 weeks shrink to about 19 days. Plugging in different values of for cases where code writing gets even faster gives the following.
| Code writing speedup | Code writing time | Total lead time | Overall speedup |
|---|---|---|---|
| 1x (before AI) | 4 days | 21 days | 1.00 |
| 2x | 2 days | 19 days | 1.11 |
| 5x | 0.8 days | 17.8 days | 1.18 |
| 10x | 0.4 days | 17.4 days | 1.21 |
| Infinite | 0 days | 17 days | 1.24 |
Notice that even if AI writes code instantly and writing time drops to zero, lead time only goes from 21 days to 17, a reduction of about 19%.
In other words, even if code writing became completely free, this team would still take nearly 3 weeks to ship a single feature. And going from 2x to 10x only saves a measly 1.6 days, so the return on effort spent making code writing even faster keeps shrinking.
Of course, as mentioned earlier, this calculation hides one assumption: that the throughput of every process stage other than code writing stays the same. But when code writing gets faster, does the throughput of the remaining stages really stay the same?
A Faster Link Can Slow Down the Chain
The calculation above showed that even when code writing gets faster, the team’s overall lead time improves less than you’d expect. But in reality, it goes further than that: when a stage that isn’t the constraint gets faster, the whole team can actually end up slower.
In this section, I’ll split that mechanism into 2 parts. First, we’ll use Little’s law to follow, with numbers, how lead time grows when a queue builds up in front of review, and then see how the nature of AI-generated PRs makes that queue even heavier.
PRs Piling Up in Front of Review
Let’s say there’s a 5-person team. This team creates about 40 PRs a week, and its members can review about 50 a week. Since the team can review a little more than it produces, PRs don’t pile up and usually get reviewed the same week they come in.
But now suppose this team adopts AI and the speed at which it creates PRs doubles. Now 80 PRs come in every week, but the team can still only review 50. That’s because the speed at which people read and understand code doesn’t increase the way the speed at which AI writes code does. So every week, 30 more PRs pile up waiting for review.
| Week | New PRs | Reviewed | PRs waiting for review |
|---|---|---|---|
| Week 0 | 40 | 40 | 0 |
| Week 1 | 80 | 50 | 30 |
| Week 2 | 80 | 50 | 60 |
| Week 3 | 80 | 50 | 90 |
| Week 4 | 80 | 50 | 120 |
How long a single PR has to wait to get reviewed can then be calculated with Little’s law.
In the formula above, is the amount of work sitting in the system, is the rate at which the system processes work, and is the average time a single piece of work stays in the system. It was originally used to analyze bank teller lines and factory queues, but it fits pretty well anywhere there’s a queue.
That said, this formula holds when the amount coming in equals the amount going out and the queue stays steady, so in a situation like this where the queue keeps growing, it’s best used only to roughly estimate wait time by dividing the amount of work lined up in the queue by capacity.
Doing that calculation, a PR submitted around the end of week 4 already has 120 PRs lined up ahead of it waiting for review. Given that this team’s throughput drains 50 PRs a week, this PR has to wait weeks just to get reviewed. The interesting part is that before AI, no PR had to wait 2.4 weeks like this, and that the lead time for each PR to finally get through keeps growing.
No matter how much AI the team runs to write code fast, what it actually ships to users each week is only the 50 that pass review, a measly 25% more than the 40 before AI.
The speed of creating PRs went up 100%, but throughput only went up 25%, and in exchange, lead time keeps growing every week. The dashboard will show the number of submitted PRs doubling, but the strange situation arises where the time for a single feature to reach users has actually gotten longer than before the adoption.
In other words, this team got 25% faster by throughput but actually slower by lead time. As long as the queue in front of the bottleneck is full, throughput is fixed at review’s capacity of 50, and only lead time grows as the queue gets longer.
Of course, this is an example with simplified numbers. Real teams react when a queue builds up, by spending more time on review, reviewing carelessly, or simply making fewer PRs. But however they react, the structure stays the same: on a team where review is Herbie, speeding up only code writing creates a queue.
And this darn queue doesn’t just add waiting time. If review comments arrive 2 weeks after a PR goes up, the author is already in the middle of 3 or 4 other tasks, so there’s another context-switching cost on top.
In the end, to address review comments, the author has to bring back to mind why they wrote the PR that way, and if AI wrote that code, there’s often barely any context to bring back in the first place. Reviewers, too, have to load a different context into their heads each time as they jump between piled-up PRs, so the same review can easily take longer than it would without a queue.
The problem is that this queue rarely shows up in the numbers organizations usually look at. The dashboards that AI coding tools provide out of the box usually show numbers like active users, the acceptance rate of AI-suggested code, and token counts.
And the trap is that on a team like the one in the example, all of these numbers seem to be going up in a good direction.
Growing PRs and Work That Comes Back
On top of that, reality is even rougher than this example. As mentioned earlier, AI-generated PRs generally get bigger, and since the context doesn’t stay clearly in the author’s head, they’re also harder to review.
In the Faros AI study, the average size of PRs on teams that used AI heavily grew 154%. If each PR grows to 2.5x its size, the number of PRs a reviewer can get through in the same amount of time drops by that much. As I mentioned in my June post, the cost of generating code has converged to nearly zero, but the cost of reading and understanding that code hasn’t dropped at all, so the backpack Herbie, review, carries has actually gotten heavier.
On top of that, there’s also the problem that regressions happen more often than before in AI-written code. The fact that delivery stability fell as AI adoption increased in the DORA 2024 report means that deployed changes causing incidents or getting rolled back became more common.
And a feature that regresses like this goes right back to the front of the team’s chain. You have to find the cause, write a fix, get it reviewed again, and deploy it again. One more piece of work gets piled onto a review queue that’s already clogged by the bottleneck.
Put the other way around, raising the quality of later stages to reduce work that comes back is effectively giving the bottleneck more time to spend on new work, which makes it one of the few improvements behind the bottleneck that can actually raise throughput.
Google’s announcement from the introduction is worth rereading from this angle as well. That 75% of new code is generated by AI and approved by engineers also means, flipped around, that the amount of code engineers have to approve has grown by that much. For an organization like Google, where review, testing, and deployment are highly automated, that approval may be manageable, but whether the same holds for other organizations that set the same number as a target is an entirely different question.
A 2025 study by GitClear, a codebase analytics company, points in a similar direction. In 2024, code blocks duplicated across 5 or more lines increased 8x over before, and the share of refactoring-style changes that move and clean up existing code fell from about 25% in 2021 to under 10%. Pasting in new code got easier, while cleaning up existing code happened less, and this comes back not as today’s review queue but as the cost of changes a few months down the line.
In other words, in a chain where only the code-writing link got faster, a queue builds up in front of review, bigger PRs make review even slower, and unstable deployments send work back to the front of the chain. This is how the sad situation arises where every number on individual dashboards improves while the speed at which the team reaches users stays the same or may even slow down.
So Why Doesn’t Anyone Measure It?
With results this mixed, you’d think organizations would be eagerly measuring the impact of AI adoption, but as we saw earlier, most of them don’t. It seems that for most, the reason is something other than not knowing how to measure.
In this section, I’ll split the reasons organizations don’t measure into 2, and finally look at how something similar played out in factories 100 years ago.
Because Everyone Else Is Doing It
Once Google announces a number like 75%, leaders at other organizations start getting asked, “What’s our percentage?”
When no one knows what the right answer is, organizations tend to copy those who look successful, because copying is much safer for the leader personally. If everyone else is doing something and we alone don’t and fall behind, it’s the leader’s fault, but if we did the same thing everyone else did and it didn’t work, it’s a problem for the whole industry.
In this situation, the adoption rate becomes not a metric for measuring impact but a signal that shows we’re doing something too. And once the signal becomes the goal, measurement turns into something there’s no real reason to do, since the goal was achieved the moment adoption happened.
The pressure doesn’t come from just one direction, either. When headquarters, investors, or big customers start asking how much AI is being used, the organization has to raise its adoption rate regardless of the impact. And now that the idea that developers who don’t use AI are falling behind circulates daily at conferences and on LinkedIn, using AI is becoming a matter of identity for individual developers before it’s a matter of efficiency.
The sociologists Paul DiMaggio and Walter Powell called the phenomenon of organizations in the same field coming to resemble one another through imitation, coercion, and norms, regardless of efficiency, institutional isomorphism.
So it’s hard to see this as a problem caused by one or two leaders being lazy about measurement. From above, they’re asked about adoption rates; from the side, they hear competitors’ numbers; and from below, developers feel they’ve gotten faster on their own, so the person who steps up to say “let’s measure” ends up being the one going against the current. Perhaps this structure is why so many organizations don’t measure the impact of AI adoption.
Because a Prediction Written Down Can Be Wrong
Another reason is the risk that comes with measuring. In my September post, I said the numbers you’re going to measure should be written down before you act. If you look for numbers after the work is done, at least one of dozens of metrics will have improved, even by chance, so numbers picked after the fact aren’t very convincing.
But AI adoption is a case where this prescription is especially hard to follow. If you wrote down before the adoption that lead time would drop from 3 weeks to 2 and it actually stayed the same, that amounts to publicly revealing that a confidence everyone shared was wrong.
So it becomes a much safer choice for the individual to not write down a prediction at all, and to pick numbers that went up after the adoption and report those. Numbers like submitted PRs or lines of code written by AI have almost always gone up, after all.
In fact, there’s a field that blocked this problem with an institution quite a long time ago: medicine. After the practice of publishing only the metrics that came out well after a clinical trial kept repeating, the editors of major medical journals issued a joint statement and decided that starting in 2005, they wouldn’t publish clinical trials that hadn’t publicly registered what they’d count as outcomes before the trial began. They turned writing down predictions in advance into a rule rather than a matter of individual conscience.
In the same way, I think an organization’s AI adoption should also decide “what would count as success if it improved” before adopting, and record it somewhere everyone can see.
Forty Years for the Electric Motor to Change the Factory
In fact, this isn’t the first time this has happened. In a book review he wrote for the New York Times Book Review in July 1987, Nobel laureate in economics Robert Solow left this line.
You can see the computer age everywhere but in the productivity statistics.
This remark, that productivity was underwhelming compared to all the excitement over a new tool called the computer, later came to be known as the productivity paradox, and the answer the economic historian Paul David offered for this paradox was that the electric motor 100 years earlier had been exactly the same.
In the early 1880s, Edison opened the first power stations in New York and London. Yet even in 1899, electric motors accounted for less than 5% of the power driving machines in American factories, and that share only passed half in the early 1920s.
And according to David, it was also from those early 1920s that factory electrification began to have a noticeable impact on manufacturing productivity. It took 40 years after Edison’s first power station opened for the effect to start showing.
So why did it take so long for the effect to show? What David focused on was the way factories used electric motors. In the steam era, a factory put one giant steam engine in the middle of the building and distributed power to every machine through long shafts and belts running across the ceiling. So to transmit power efficiently, machines had to be crowded near the shafts, and as a result factories were built tall, with many floors.
Most of the factories that first brought in electric motors simply swapped a large motor into the steam engine’s place and, just as before, had a single motor turn the shafts to drive many machines at once. This approach was called group drive, and according to David, it was dominant from the mid-1890s until just before the 1920s. With the shafts, the belts, and the machine layout all unchanged, there was no way the factory could get faster.
Factories actually got faster only once unit drive, which put a small motor on each machine, spread. Now that machines no longer depended on a central motor, they could be freely laid out in the order of the work, factories were built wide on a single floor, and the paths materials traveled got shorter. David argued that about half of the sharp rise in American manufacturing productivity growth between 1919 and 1929, compared with the previous decade, could be explained by this increase in individual motors.
In other words, what changed productivity wasn’t the new tool, the electric motor, itself but redesigning the entire chain of the factory around that tool. I think the way many organizations use AI today looks a lot like the factories of the 1900s that swapped a motor into the steam engine’s place. They slotted AI into the place where code gets written but left the chain of planning, review, QA, and deployment as it was.
Redesigning the Chain
A hint about what needs to change can again be found in Goldratt. He laid out a way to actually apply the theory of constraints in 5 steps: identify the constraint, exploit the constraint as much as possible, subordinate everything else to the constraint, elevate the constraint itself, and repeat. In this section, I’ll carry this sequence over to software teams and lay out what needs to change for a team that has adopted AI to actually get faster.
To say it up front, what I talk about here is closer to a starting point than an answer. Each team’s Herbie is different, and even within the same team, Herbie moves around over time.
That’s also why the 5th step is to repeat. Goldratt stressed that once you resolve one constraint, something else always becomes the new constraint, so you shouldn’t let the inertia of fixing the old constraint become a new constraint. And in this section, borrowing from the Boy Scout story, I’ll keep calling the team’s slowest link Herbie.
First, Find the Slowest Link
The first thing to do isn’t to decide where to use AI but to find where your team’s Herbie is. And the surest way to find it is to write down, separately for each stage, the time spent working and the time spent waiting as a single feature goes from request to deployment. (In the end, measurement matters most.)
As in the September post’s example, once you see that code writing takes 4 out of 3 weeks, you just need to chase down where the other 17 days went. If review waits are long, review is Herbie; if waiting on another team’s API takes a long time, dependencies between teams are Herbie; and if deciding what to build takes a long time, planning and decision-making are Herbie. For quite a few teams, Herbie is probably more likely to be before or after code writing than in it.
You don’t need fancy tools, either. Quite a few teams already have the data they need: all it takes is the time a PR was opened, the time the first review arrived, the time it was merged, and the time it was deployed. And these days, you can whip up the data collection and analysis quickly with MCP, so it isn’t that hard.
The 4 timestamps collected this way leave 3 gaps between them: the time spent waiting for review, the time spent going back and forth on review, and the time spent waiting for deployment. Add the time a request came in and the time development started from your issue tracker, and you can even see the time that passed before code writing.
Gather a few weeks’ worth of PRs this way and take the median for each gap, and you can roughly see where your team’s 3 weeks are going. It doesn’t need to be precise. As I said in the September post, measurement isn’t about proving something but about making what you don’t know a little less unknown, and at this stage, all you want to know is roughly where Herbie is. (If you get greedy about precision here, you’ll find yourself in an endless grind, so be careful.)
It’s best to do this before adopting AI, because that way you can also see where Herbie moved after the adoption. As in the earlier example, when code writing gets faster, Herbie tends to move to review, and when you speed up review, it moves again to QA or deployment.
The Temptation to Cut the Slowest Link
But once it becomes clear that review is the slowest link, these days a rather different kind of talk comes up as well: “In the age of AI, do we really need reviews where humans read code line by line?” The logic goes that AI writes the code and AI adds the tests, yet humans are reading all of it, which creates a bottleneck, and so I hear quite a lot of talk around me about getting rid of the review process altogether or swapping it out wholesale for AI review.
Whenever I hear this, I think of Herbie. If the line is running late because of Herbie and you leave Herbie behind in the woods, the line might look faster, but that isn’t the line arriving. The goal of the hike isn’t for the kids in front to arrive first but for the whole line to reach the destination. Removing the slowest link is less like shortening the chain than like breaking it.
What bothers me more about this question is how the decision gets made.
I hear a lot of talk about getting rid of reviews since AI will read the code anyway, but it seems hardly anyone says this after measuring how much the quality of AI-written code has actually changed, or what and how much review is catching right now. The decision to get rid of reviews can end up being made the same way AI adoption was, without measuring its impact.
If anything, the signals we saw earlier, such as DORA’s drop in delivery stability, Faros AI’s 9% increase in bugs per developer, and GitClear’s 8x rise in duplicated code, indicate that there’s more for review to catch. And once you get rid of review, what review was blocking is likely to surface not at the review stage but as incidents, rollbacks, and the cost of changes months later.
What’s more, review isn’t only about finding bugs. In a study of developers at Microsoft, finding defects was the biggest motivation for review, but actual reviews turned out to be less about defects than expected and instead produced benefits like knowledge transfer and team members knowing about each other’s work.
A study that analyzed code review at Google also found that what developers expect from review centers less on finding defects than on education, teaching and learning from one another, and on making code readable and understandable. As with the McNamara fallacy I mentioned in the September post, once you start treating hard-to-measure effects as if they didn’t exist, review easily looks like a process that’s nothing but cost.
Of course, it’s true that review has gotten harder. PRs have gotten bigger, and there are more cases where even the person who submitted the code has trouble explaining why it was written that way, so both the amount reviewers have to read and the context they need to make sense of it have become much heavier than before. Still, I can’t help thinking that getting rid of the process itself is less about solving the problem and closer to a rationalization for avoiding that difficulty.
As I said in my June post, the moment you hand even review over to AI, the causality behind that code no longer remains in anyone’s head on the team. Making review faster and skipping review are different things, and all the steps that follow are about making review faster.
Lighten the Load on the Slowest Link
Once you’ve found Herbie, you need to think about how to use AI on Herbie’s side. In the Boy Scout story, what made the line faster wasn’t making the fast kids walk faster but lightening Herbie’s backpack.
On a team where review is Herbie, it’s better to use AI not to write more code but to make the code under review smaller and easier to read. That means splitting big changes into reviewable small units, stating the intent of a change clearly in the PR description, and adding tests so reviewers don’t have to run the behavior through their heads step by step. I think it’s in the same vein that the DORA 2025 report named working in small batches as an organizational capability that amplifies the effects of AI.
There are also many teams whose Herbie isn’t review but comes before it. For a team that takes weeks to decide what to build, AI has a bigger impact when it’s used to make decisions faster rather than to write code faster. For example, instead of holding meeting after meeting over a planning document, you build a few simple prototypes with AI, try them out hands-on, and pick one.
Verbal debates rarely reach an end, but put 2 working screens side by side and you often reach a conclusion surprisingly quickly. Now that the cost of building a prototype has dropped to nearly zero, the cost of putting off a decision has become relatively much more expensive.
There are also teams where QA is the slowest link. On a team where finished features line up for days waiting for a manual QA slot, writing code faster only lengthens the queue in front of QA.
On such a team, it’s better to use AI to automate repetitive regression tests, or to generate draft QA scenarios from planning documents and change details so that QA staff can start from reviewing and refining a draft instead of a blank document. Judging which scenarios matter is still a human’s job, but just changing the starting point can make the load quite a bit lighter.
Pace the Other Links to the Slowest One
The third step is a bit counterintuitive: match the pace of the stages that aren’t Herbie to Herbie. In the earlier example, if the team can review 50 PRs a week, making 80 a week only leaves 30 piling up in the queue. It may be faster for the team as a whole to match the pace of creating PRs to the amount that can be reviewed and spend the leftover time helping with review or on work that doesn’t need review.
One way to actually do this is to limit work in progress. It’s a rule like: once the number of PRs waiting for review passes a certain count, instead of opening a new PR, you review someone else’s PR first. By Little’s law, when capacity stays the same, reducing the amount of work piled up reduces lead time by that much.
Of course, if the queue goes completely empty, the bottleneck ends up sitting idle and throughput drops. So the goal isn’t to eliminate the queue but to keep just a small queue, enough that the bottleneck never sits idle. Goldratt framed this as drum, buffer, rope: the drum is the pace of the bottleneck, the buffer is the small queue left in front of the bottleneck, and the rope is what ties down the earlier processes so they can’t push in more work than that. The rule about the number of PRs waiting for review mentioned above plays exactly the role of this rope.
That’s easier said than done, because individual metrics look bad. A developer who makes fewer PRs and reviews other people’s looks like the slowest person on a dashboard that only shows the number of PRs each person submitted. So at this step, what metrics the leader looks at matters more than individual resolve.
Grow the Slowest Link Itself
If the previous three steps were about making the most of the Herbie you have now, the next is to increase Herbie’s own capacity. In terms of the earlier example, it means raising the number of PRs the team can review in a week from 50 to 75.
When you look closely at a team where review is Herbie, review is often concentrated on a few people. Only one or two people really know a particular module, so every PR touching that module lines up in front of them.
On such a team, the way to increase how much can be reviewed isn’t to make reviewers work faster but to increase the number of people who can review that code. Documenting domain knowledge, bringing less experienced team members into reviews, and cleaning up module boundaries so that no one person has to know everything all fall under this.
This is closer to an investment than something done for immediate speed, and it takes quite a while for the effect to show. But I think the apprenticeship problem I talked about in my June post is actually tied to the same issue. Having less experienced developers take part in review increases review capacity, and at the same time, it’s one of the few places where those developers can build a sense for judging code in an age where AI does the code writing.
AI can help with this process, too. AI is pretty good at finding where a change has impact when a reviewer is reviewing changes to a module they’ve never seen, or at summarizing related past change history. But as I said earlier, the judgment still has to be made by people, and AI should stay on the side of helping reviewers reach a state where they can make that judgment faster.
Meanwhile, on a small team, the person doing review often handles QA or deployment as well. On such a team, the time freed up by automating later stages becomes time that can go straight into review, so improving later stages can itself become a way to grow the bottleneck.
And once review gets fast enough, Herbie moves somewhere else again. This time it might be QA, it might be deployment, or it might be back at the beginning, at the stage of deciding what to build. Just as Goldratt made the 5th step to repeat, finding Herbie is less a one-time task than a loop you have to keep running.
Leave Something to Compare Against
The last step is measurement. As we saw earlier, the reason organizations avoid measuring seems to be less that they don’t know how and more that measuring is uncomfortable. So rather than leaving measurement as a matter of willpower, I think it’s better to build it into the structure from the start.
One part is writing down predictions before adopting. Write down, even roughly, how much you expect lead time to drop and what will happen to delivery stability, and check them against reality each quarter. If a prediction turns out wrong, that’s not a failure but a signal that you had the wrong idea about where your team’s Herbie was.
The other part is leaving something to compare against. Adopt first in one team or part of the work, and put lead time and stability side by side against the part that didn’t adopt. Like the DORA metrics I mentioned in the September post, you need to look at speed metrics such as deployment frequency and lead time together with stability metrics such as change failure rate and time to restore, because as we saw earlier, stability is the first thing AI shakes.
That’s also why the metrics at Toss that I mentioned in the introduction try to measure not only productivity but application stability as well. If you only look at the numbers that say things got faster, it’s easy to miss the side that regressed.
If it were me, I’d start by looking at just three: the median lead time from request to deployment, the median time PRs waited for review, and the share of deployed changes that led to a rollback or a hotfix.
All three are rough numbers swayed by factors other than AI, but as I said in the September post, checking whether a few imperfect metrics point in the same direction beats measuring nothing while searching for one perfect metric. And only if these three are improving solely on the side that adopted AI can you say that AI is making the team faster.
Of course, even this won’t cleanly prove the impact of AI, and in the end, you’ll have to somehow make a judgment holding a few rough numbers. But at the very least, it can keep you from believing the team got faster just because the number of submitted PRs went up.
Closing Thoughts
This post started from the question “Does just adopting AI really make a team faster?” And as I kept researching, I came to the conclusion that AI, at least so far, only speeds up one link in a team’s chain, and that a team’s speed is set not by that link but by the whole chain. On a team where code writing wasn’t Herbie, speeding up only code writing makes the team faster by less than expected, or even slower as the queue piles up.
But what stayed with me more was the next question. Why do so many organizations not check this, even when the results are this mixed? Because the fact that we’re doing what everyone else is doing has already become the goal, and having a prediction you wrote down in advance turn out wrong is far too uncomfortable for an individual. In the end, the question of whether adopting AI makes a team faster is closer to the question of whether your team knows where the weakest point in its own chain is than to a question about AI. And the organizations that haven’t answered that question are the ones most likely to lean toward breaking the clogged link off instead of looking into it.
The reason it took 40 years for factories 100 years ago to actually benefit from electric motors wasn’t a lack of motors but that they swapped a motor into the steam engine’s place and left the factory as it was. And Solow and David could point out this paradox in the first place because someone had been steadily measuring the output leaving the factory gate. If no one had measured it, nobody would even have known that the factories that swapped in motors had stayed the same.
Software teams are no different. The number of submitted PRs and the lines of code written by AI only show how busy the machines were, and whether the team actually got faster can only be known by measuring at the exit. You have to measure how many features were actually deployed to users, and how long it took for a single feature to reach them. That’s why I think what you need to decide before where to slot AI in is where your team’s exit is and what you’ll measure there. Only once you start measuring that does it become clear where your team’s Herbie is.
And with that, I’ll bring this post on whether adopting AI really makes teams faster to a close.
관련 포스팅 보러가기
What Leaders Should Really Worry About Isn't Productivity
Essay/CareerOn Properly Defining Ambiguous Problems
Essay/CareerWhy Can't Claude Compute the Weather?
Programming/Machine LearningDevelopers Who Stopped Growing
EssayOnly When the Tide Goes Out Do You See Who's Been Swimming Naked
essay