# Measuring AI ROI in Software Engineering

**Podcast:** Dev Interrupted
**Published:** 2026-08-11

## Transcript

Welcome back to Dev Interrupted, brought to you by Linear B.
My guest today is Yishai Beery, Linear B's CTO, and he spent the last few months exploring at Linear B whether this AI spend in the industry is turning into measured, delivered work.
Because something shifted in the last year for engineering teams, and we're all still grappling with it.
Some teams pulled ahead with AI, others stayed flat, and some even took a dive.
And left unchecked, that gap can rip your engineering team in half.
So in today's episode, Ben and I dig into how to close that gap, improve AI ROI.
Here's our conversation with Yishai.
AI has changed how software gets built faster than any single year of our benchmarks could possibly capture.
So this month, Linear B is publishing a mid-year refresh, built on a fresh data set of millions of pull requests from hundreds of engineering organizations from February to May 2026.
We've been digging really into AI data within this data set, you know, AI usage where it's being adopted and leveraged.
And the results have been striking and very clear, as a matter of fact.
And the fact is that there's a gap that's opening between teams that are getting significant leverage out of AI and the teams that have merely just switched on or maybe not even really adopted it yet.
And that gap is widening at an extremely rapid pace.
So we wanted to bring our guest on today.
Because he has a very clear picture of all of this data that we've been collecting and what it means for engineering leaders.
So I want to welcome Linear B's CTO, Yishai Biri.
Yishai, welcome to Dev Interrupted.
Thank you.
Yeah, great being here again.
Great to have you here.
Yeah, so what we're really focusing on is, you know, how there's been this significant shift that we have now seen starting in January of this year.
And everyone has started to pay attention to like AI costs and like what they're getting out of it.
And we're in this critical moment right now, like at this point in time where we are beginning to see teams that are effectively investing into AI and getting significant improvements, productivity improvements out of it.
So we normally do our sort of engineering benchmarks once a year, sort of at the start of the year to kick everything off.
But we felt that things were changing so fast that we had to start covering this on a more regular cadence.
So, Yishat, I just want to start there.
Like, if we weren't looking at this data here in the middle of 2026, what do you think we would be missing?
Like, what are we getting out of doing this refresh of our benchmarks data halfway through the year?
I think two things at play here.
One is behaviors are beginning to change rapidly across the developer organizations that we're looking at.
waiting a full year to get a fresh grasp of what are the behaviors, how do the common metrics look like, that's too slow.
Changes are so fast, you want to make sure that your benchmarks are actually capturing what's going on now and not what happened six or nine months ago.
The other thing is a focus on some new benchmarks around AI, how it's getting used, the cost, etc.
The focus in the market around these questions has really increased.
And, you know, if a year ago or even six months ago, people were asking about adoption, now it's about cost and ROI.
This has quickly become the number one priority.
I'm speaking with, you know, dozens of engineer leaders and everyone is saying the costs are growing and we're still, they're not finished growing.
We all know that.
I need to show actual ROI.
This is now CFO, CEO level visibility.
I need to show that this token spend is actually giving back hard business value and it's worth it.
It's worth the hundreds and sometimes thousands of dollars a month for a developer.
That focus means that we have to be talking and showing metrics and benchmarks and data about how AI is being used and is it getting the or creating the right leverage, the right...
productivity benefits to actually justify all of this cost.
And then there was this, there's always been hype, right?
People reading about my dev team is 2x, 5x, 10x companies saying I can cut half of my workforce because of AI.
All of this hype now has to meet reality.
So mapping, are we there yet?
Is the hype actual, like, is it real?
What are the gaps and how is that meeting a specific developer organization or?
The market at large, that's what we set out to do here.
And this cannot wait to an annual cadence of refreshing benchmarks.
Yeah, I think it's really smart to point out that if you were to wait a year, you'd be missing so many learnings, patterns, behaviors, and honestly stumbling for that year.
Because the reality is, is that the adoption of these tools, the impacts of them are in more real time.
than they've ever been.
These rollouts are more aggressive.
Folks are building more things more quickly.
And just the need for more real-time awareness over what these tools are doing in your organization and what you're getting out of them is just, it's never been higher.
And it really does tie to the whole story of cost, like you were saying a moment ago, because Like we're now we're working with our technology almost on the same kind of cadence as how financial leaders think about how they're going to be thinking of their budgets.
Like on a quarterly basis, how are we justifying?
What did we do last quarter?
What are we doing this quarter?
Now our tools and our AI spend are.
right in those conversations.
And so we need just as much real-time data that justify the things that we're doing with those tools and understand the downstream impact.
And like finance, by the way, is like really in the room on these adoption conversations.
You know, folks that are leading these budgets at larger organizations are really, really obsessed with the ROI question.
They're coming to their technology leaders, many of whom are listening to our show, and they ask them like, prove it, prove your AI rollout is working.
And that's where, you know, folks turn to things like metrics and how their tools are getting adopted to understand how it's moving through their org.
And maybe you could shine some light on us for this ROI question.
And it's so predominant.
Like you mentioned it even a moment ago that it's all around cost.
You know, why has this become such a number one lens?
And, you know, why does that push us to benchmark what productivity means?
How does that help that conversation?
So I think finance is not just in the room.
They came in late.
They kicked the door in, frantically trying to get on top of what's happening.
I've heard this from more than one of our customers where, you know, first, let's just go all in on AI.
Token maxing, experiment.
We can't miss the boat.
Have everyone use it.
Let's see what happens.
And let's be pretty honest.
All of us are still experimenting.
A lot of it is maturing, but...
There's no mature model for using AI in a dev work yet.
It's all very, very experimental.
Everyone has to play, has to try the new shiny things because the impact can be so great.
And now is the cycle where finance is starting to see, okay, the budget we had placed, like it was earmarked because no one knew how much is it actually going to cost.
And these budgets were blown away, you know, in two or three months, the whole budget for the year.
There was a management offsite and we came back with a clear mandate.
We have to show ROI for the spend.
Not because we're going to cut it tomorrow, but because it's not large enough to show up on the CFO's radar.
It's no longer a side thing.
It's becoming, I don't know, 5% of my dev spend, 10% of my dev spend, maybe 20% in some places, which is substantial.
It's now becoming a real thing.
And combine that with the price hikes.
And, you know, there was a lot of recent activity in the last three months with models exposing more of the actual cost to customers.
I think we all know we're still heavily subsidized when doing like coding with AI in most of the footprints.
So there's still room to grow for those prices to grow.
We're not paying the actual inference cost yet.
And that like dawned on those organizations.
We have to start looking at not just spending whatever we can to get on top of the technology, but also how do I start introducing some kind of structure and process to my spend decisions and my spend budgets?
I think these just culminated now in a very sharp focus on how much is it actually costing me?
Can I attribute these costs to specific projects or types of work or teams and so on?
And I don't know, the golden, the holy grail, what is the ROI?
What is it giving me back?
Even when all my developers are saying, this is great, I need something more concrete.
I need to show this is actually helping me move so much faster.
I feel like the story of last year was like, Are people actually using AI, like as you mentioned?
And I feel like that question was sort of like put to rest basically at the end of last year.
It's like every study that was looking at this, including our own data, when we looked into this was saying that like basically 90 plus percent of developers have adopted AI in some manner within their daily or weekly workflow.
And, you know, so that's like no longer the thing that people...
question anymore.
It's like it just sort of assumes that AI is now ubiquitous across the organization.
And it's not really a thing that actually indicates whether or not you're seeing productivity gains out of it.
Really, what it comes down to is how much leverage you're getting from it, like how much how much of AI's outputs are actually showing up and things that provide value to your customers, you know, or that get shipped into production.
We've been sort of cautious at Linear B to like fully lean into like just measuring adoption because we understand that it's very limited in terms of like what it can actually tell you.
Even at the most like AI forward companies.
Explain this like leverage challenge to me.
Like why is it, you know, now that we've moved beyond adoption metrics, like what is it now that we need to be looking at to really understand where AI is impacting us?
Yeah.
So when you assume, and that's typically a correct assumption that Your developers are using AI in some shape in their daily work or in their work, typically as coders or people bringing new code to the code base and creating value to the company through that.
You're asking, okay, what is it?
What are the returns?
And there's many, many factors that will influence what I'm actually getting from AI or from that new tool that I'm paying for in the hands of my developers.
So this could be about, yeah, the same task.
I can write the code much faster or Claude or whatever AI tool I use can write the code for me much faster.
But coding is just a small part of what developers do.
And coding is not enough to deliver value.
So you have to look at how is AI helping me move faster?
And there is a plethora of behaviors and changes in how the SDLC works with AI that will...
impact or, you know, limit or maybe unblock the kind of value I'm getting from AI.
Let me give you an example.
If AI helps me write code faster by that code is sloppy and I will need to rework it in two weeks because of runtime bugs or because of, you know, production failures, then I did not really improve my velocity, right?
I'm not delivering that much more because now I'm stuck fixing those issues with much more expensive cycles.
The impact on quality, on the ability to make the second or third change in that code base, all of these compound into the overall productivity question.
Maybe I can get the code written fast, but I can't review it fast enough.
I have a new bottleneck with reviews and acceptance that will limit the productivity gains I can get from AI.
So yeah, I code 10x time faster, but I'm not delivering.
And certainly the team is not delivering 10x more if they're stuck on a new bottleneck.
Not to mention some softer things like, are my seniors getting stressed out by having to just review crazy AI PRs all day?
Like, are they burning out on a new type of load and new type of focus in their work?
Because I have to get, you know, people have to review code, at least in some places.
And that code is now written by AI.
which is harder to review.
It's typically larger, but the PRs are larger.
A lot of factors, and then that becomes a problem with my talent over time.
And my growing seniors, they got like, well, next year will I have seniors?
What is the path from junior to senior in an AI-dominant SDLC?
So there's this either directly impacting productivity items or softer ones, which will take a little more time to compound.
But almost like a back to basics, there's many moving parts in developer productivity, and all of them will affect the kind of value I can get from AI.
So I may be investing in tokens or getting my people the best tools, but if I'm not solving for these problems, that will diminish or even erase those productivity gains I can get from AI.
And when we talk about leverage, it's basically saying, How efficient are you in translating the AI investment and all the goodness that AI and the SOC can give me into actual productivity gains?
Yeah.
And like you hinted at the beginning, we're seeing a very big divide in that kind of productivity gains that different teams and different organizations and even different developers get.
And I think that's where it becomes interesting.
What sets apart the people or the organizations that are able to get those 2x, 3x more productivity gains compared to the ones that are, okay, I'm getting 10%.
And what should I change to get closer to that high end of leverage?
Yeah, and I think this is a good place to just inject some of the specific data from this report because I think it's really relevant to this moment.
And we've kind of buried the lead a little bit on this, but I think everything we've discussed at this point is really important to context for it.
But what we've seen is that developers who are in the highest category of AI usage, and the way we came up with that number is if they're in the P90s, the 90th percentile of developers who are using AI for code that goes into a pull request, it's how we're looking at this.
Year over year, they've more than doubled their output.
in terms of code that has been merged into production versus people who really don't use AI for the code that goes into their PRs.
They're basically flat year over year.
Today looks just like it did a year ago.
And that's that gap that we're talking about.
And it correlates pretty well all the way down the cohort.
So P75 is also seeing an uplift, but not nearly as much as that top tier group.
And that's really sort of the headline takeaway that we got from this report is that that is what has changed since January.
These higher groups, these higher cohorts of AI users that are using it for productive work have more than doubled their velocity over the past year, which is a very significant change.
Yeah.
So I think several things here.
One is, I think, a renewed focus on...
a core metric, which is how many PRs can I get merged and normalize that by the number of developers.
So a typical developer on my team, how many PRs can get it?
Can they merge every week?
So that helps.
It's obviously like every metric is going to be a simplification of reality, but that capture is a good sense of the productivity.
And if that number grows over time, this means A, I have unblocked.
barriers to getting PRs merged.
And that's like code delivered to my code base.
And that's the business value behind that.
And the growth there means I was able to translate AI to not just writing more code, but to actually delivering more code.
There's many ways you can argue about why count PRs.
And I can split my PRs to get a better count.
I can make my PR smaller to get a better count.
All good things.
I would urge you to do that.
And I'm fine with that tack on the gap between model and reality in this metric, because the ways to cheat this metric are actually ways to win.
But maybe even more than cycle time, which for many years was the staple of our reproductive short cycles, what we're seeing with AI is typically same cycles, but more than parallel.
So with AI now being in the focus, Approach here is not to try and look at are my cycles much shorter.
In reality, they're not much shorter.
But rather, am I able to deliver more?
And that happens through more parallel production by the developers as they're using AI to get more things done.
So PR is merged by developer, like per developer per week.
That's the metric we are looking into.
Like you said, if we are comparing developers to a year back, we're seeing a wide.
range of multiples from the ones that have not changed anything and they're not using AI.
So that's a good kind of control group for us.
Yeah, nothing changed.
You're not using AI really.
And I'm not surprised that nothing changed for you.
And then the ones at the top, which are able to do two or two, even two and a half X of their previous throughput.
So that kind of gives an idea of the range.
It's not the maximum possible, but if you are A typical developer organization, we typically look at larger ones with hundreds of developers or more.
And you fall anywhere there, you're in the normal range.
And now your question is, am I getting 10%, 20%, 30% more than I used to do a year ago?
I'm in the low ranges of AI leverage, and I should be looking to get more.
If I'm in the 2x or 2.5x, I'm pretty much at the top.
I can always get better, but I'm in a good place.
And when we look at the graph, and we will obviously be sharing that, you know, the full data, you see a big jump in January onwards.
So from last June to January, things are pretty much stable with a low range of leverage numbers, if you like, or increase in PR throughput.
And then in January, things start to change.
And now the leaders begin to grow, you know, break away from the others.
I think it's obvious that there were like key model or frontier model drops that made this a reality, along with ongoing improvements in the harnesses and the tooling around this.
So for the industry, something clicked around January of this year.
And that something is now finally translating into actual productivity gains.
It's not the 5x hype.
And even 2x is a broad, like it's not the average.
2x is good.
But 2x is now achievable.
And the devs that get it, the teams that get it, and the ones that are able to remove the other bottlenecks are now getting that 2x, 2.5x, etc.
I think it's really smart that you called out that adoption hides things.
You know, if you just look at adoption, you're going to miss all these other patterns.
And behaviors, especially even going back to what you were saying about rework, like if you're just tracking token spend, you're just tracking the usage of the tokens, then it's completely hiding how much rework, refactor, fixing and cleaning up of messes that AI is making is actually happening.
And in fact, it makes it look like more good things are happening.
So it misdirects the narrative.
And also there's a really fascinating thing that happened.
Right around the end of last year, which you touched on, where folks went on, you know, they went on holiday break.
They had some time to themselves.
The models were getting kind of good.
And then they came back at the top of the year and maybe they had a hobby project that they did and they shared or they otherwise were able to experiment without just having the week to week race of getting their PRs done.
And they were able to do a reset.
at the top of the year.
Oh, this is how I work.
This is how I structure the beginning of my tasks.
This is how I do multiple things in parallel, which is, I think, the key to getting more of these impacts to tracking like, oh, these folks that are getting 2x, 3x, however many x more out of their output.
That's the real secret to them.
It's not that they're taking shortcuts or that all of their windows are super, super small.
It's that they've stacked a bunch of those windows on top of each other.
And managing that context and how you set up all of that work, get it going during the day, check back in on different things and then close it all out.
That's a whole new style of working that engineers are still getting.
comfortable with.
And that translates itself into things like metrics and understanding like our PR merge rate and how big are these PRs?
Like it shows up in a lot of different places.
I also really love that you.
You called out that in this case of, yeah, you could totally game this metric, but in gaming it, you're actually probably doing everyone a favor because you're writing smaller, cleaner, more atomic PRs, and you're just working in probably a more better fashion for your teammates to review.
Another thing, too, that I think is just so, so critical, and I got to just touch on it again, is the idea of avoiding the rework problem.
Being really aligned up front now is the biggest task at hand.
Talking with your peers, talking with your stakeholders, talking with your customers, getting really aligned about what it is you need to fix first before you start throwing tokens and people at it.
This is what really allows folks to...
stack these windows on top of each other.
They can trust that they did their homework, that folks know what they want, that the agent has the tools they need, and they can come in at the end and review and provide the quality gate.
And this is what allows folks to get a lot of that.
That's what really I think is accelerating the widening of that gap.
I know myself as someone who runs a lot of sessions and works with agents all day when I talk with folks who don't have the parallel.
parallelizing action, that gap really is evident in conversations even.
And so going back to what's driving all of this, you know, and that is the token spend.
I'm really excited that we refresh these benchmarks because we get to reflect all of these new changes and everyone's working a pattern since the top of the year, which as you called out are pretty stark compared to the end of last.
And, but now we're starting to get really early data on how like folks are putting a number on their token spend.
And, you know, Customers that we're talking with, they're watching that bill climb in a way that they can't really necessarily feel like they have a grapple on.
And it gets really hard to tie that to results.
So, you know, I'm kind of curious, like, obviously, the anxiety around that conversation is kind of everywhere.
When you go and you talk with folks about how are your teams using AI?
How much is it costing?
What are you getting out of it?
What stands out for you for companies where you talk with them?
And, you know, maybe the rising AI budget, it stops being a worry.
becomes the whole point, like they're fully leveraged on it.
What do you think really stands them apart in conversations?
One common thread that I'm hearing is, yeah, budget is a, or cost is a concern.
We have to get better at measuring it.
We have to understand our why.
That's the first like ask from, from senior management.
It's not stop.
It's too much.
It's more like show me the value.
So being able to connect the dots.
and show the value, show the increase in productivity connected to that spend, that becomes the first priority.
Then companies are starting to put in some guardrails, typically in the form of, okay, there's going to be a monthly limit for a seat or for a developer, typically a generous one.
So more than my budget, but at least it makes sure nothing runs astray.
And I've heard of companies using a $2,000 limit or a $1,000 limit per month.
Again, Not always reaching that limit, but it's kind of like a guardrail for mistakes, which can be expensive.
And I imagine that in the next few months, as more visibility around cost and ROI bubbles up and is available to those companies, now there will be like a fine-tuning cycle where they're saying, okay, if I'm getting great leverage, by all means, let's spend more.
Every dollar I put into the tokens.
If my developers are blocked, let me unblock them, give them more budget because they are showing me that this translates into many dollars in value in returns.
It's a high ROI investment.
By all means, let's do it.
If my leverage is poor, I'm only get, I don't know, if I'm spending 10% of my dev cost on tokens and I'm getting a 10% increase, that's not a great ROI.
Maybe it's positive, but I have other investments that are better.
So it's going to be kind of a pendulum between this team or this part of my organization gets it.
I'm going to give them more.
This team still needs to learn and unblock.
I'm going to invest in unblocking before I spend more tokens there.
And think about it almost like marketing campaigns, right?
I want to be cynical for a minute.
If this campaign gives me, you know, a lower CPL or more bang for the buck, it's going to get more dollars.
The other ones are going to be switched.
Either unblock teams that are already getting it and showing high leverage, take more dollars, use AI more because you're already multiplying it correctly.
Other teams, go learn how to improve from the better teams, from the teams that already unblocked themselves.
Make sure that you've got the quality in place, that you have the review and acceptance bottleneck solved, that you were investing in the right places.
Maybe coding is not your bottleneck and it never was.
There are teams that spend a lot of their time fixing bugs in production, right?
That's the reality of their product.
So triage, getting from first report to a fix quickly.
There's a lot of places where AI can be used.
It's not always just coding.
Analyze the logs, whatever.
There's a whole lot.
Not to mention documentation.
There's so many things besides coding.
that live in the SDLC and AI can really solve for.
But until you show me you have great ROI on your investment in good AI leverage, I'm going to give you a modest budget to work with and then unblock you when you were showing the actual returns.
And blend into all of that, the token budget is now almost becoming like a perk or a job requirement, right?
People...
There are engineers who will not go to work for a place that has a very low budget that gives them only so-and-so tokens and they can't run.
So all these considerations live together.
But in the end, this is about almost like a new kind of, a new way of managing the dev organization in terms of the finance.
It used to be headcount.
That's the only cost.
That's the main cost.
Get good talent.
They can do the work for you.
There's a growing percentage of the budget that's not headcount, it's tokens.
And if it's 10% today or 5% or 10% today, it's going to be 20% or 30% tomorrow.
That's no longer something you can ignore or, you know, even consider secondary.
So now, how do I balance these things together to get the right kind of impact I need?
Your SDLC looks more like a software factory every day.
How do you get ahead of that transformation?
And how do you prove what it cost and what it delivered?
On August 27th, DevInterrupted hosts a live roundtable on this very topic, proving AI ROI from software factories.
To learn, we've invited two industry experts and past DevInterrupted guests.
It's Dex Horthy of Human Layer, who ran a fully automated factory and then shut it down.
And Zach Lloyd of Warp, who publishes frequently about how he measures what his factory pays for itself.
Linear B co-founder Dan Lyons will join them to discuss the power of the context layer that will make all of this possible.
Save your seat on Luma.
Yeah, and you brought up some really great points.
One of the topics that I think has emerged from this is this concept of yield rate, right?
Like, what are you actually yielding in terms of pull request output from, you know, whether it's a human generating it or AI fully generating it or human and AI working together?
Or even in some cases of an autonomous agent or a semi-autonomous agent.
But one thing we see in this data, there is a measurable decline in the yield of pull requests the more that an AI is involved with it.
To the point where fully autonomous coding agents, despite all of the hype that's out there, they don't really seem to be generating a significant impact on productive output.
So, you know, if we have AI helping us generate all of these pull requests, and we are seeing that some teams are doubling, you know, they're 2xing their output today, but other teams still aren't, you know, what's happening with all this AI usage?
Like, why are things getting eaten up in the process despite all of this additional output?
I think there's a good reason why pull requests are still the main interface for handing off work from a...
creation phase to an acceptance phase.
Even when AI is involved, even when things are agentic, PRs represent a good, clean interface, a good cutoff.
This is where the tests will eventually run.
External tooling, reviews, all of these processes typically run on PRs.
And I think that's going to stay like that for a while.
It represents a very solid representation of This is the change that someone is proposing, and this is all the metadata attached to it, all the results of tests.
It's a very clean interface.
So when looking at PRs and the PR yield question, this is about, okay, there's a bunch of incoming work in the form of incoming PRs.
People, agents, harnesses, they all eventually create pull requests, which represents a diff I want to push into the code base.
So that is incoming work.
Not all of that gets merged, right?
Some of that gets rejected, ignored, eventually abandoned.
And the ratio here is the yield.
So if I'm creating 100 PRs, but only 90 get merged and the rest just remain unmerged or get closed, rejected, that's a 90% yield.
And what we're seeing in the data that when humans are in charge, when humans are creating code, the numbers...
like the yield rates are pretty high.
They can be 90% down to 85, maybe low 80s.
That's a typical merge rate for human work.
When AI is involved, if it's just AI help me coding, it's still a human in charge.
And the merge rates are going to be only slightly lower.
So two or three percentage points lower.
And we can attribute that to the difficulty in reviewing.
AI codes, the larger PRs that you typically get from AI, and so on.
When you look at the fully agentic flows, where an agent lives in a loop, pulls something off of a Jira queue, implements something, pushes a PR, and you've detached the human ownership part, you now see a dramatic decrease in yield rates all the way down to 30%.
So the agent creates three PRs, only one will actually make it.
to the code base.
And I think the obvious attribution here is to lack of ownership because a lot of what, you know, developers in a team need to do is not just write the code, but actually push it, make sure it gets the right review, make sure it gets the approvals, passes the test, everything that's needed.
And it's not enough that it passes.
You actually have to chase people to get the approvals and to actually merge it.
So humans pushing their work and Part of owning it is to get it done.
When that's not happening, you will now have a queue of PRs that no one's job it is to push it and get those merged.
Maybe I'm going to review them.
Okay.
But if I have a comment, someone has to push it through.
Someone has to chase people.
That's not happening with agentic PRs, like fully agentic.
And I think that plus the fact that agentic loops are still very much experimental.
So to begin with, they're not creating a lot of the volume of PRs in the large commercial developer organizations we're talking about in the real world.
There's obviously pockets of higher adoption, but it's still experimental.
And then in many cases, that loop lives on low priority items.
So it's taking the low priority bugs or fixes or feature requests that no one cared about enough to begin with.
And now it's creating PRs that no one cares enough to actually chase and push forward.
That altogether gives you a vanishingly low yield rate.
And until the ownership problem is solved, I think these kinds of flows are not going to make it into a dramatic impact on my overall delivery as a team.
So there's still some solving to do there.
But Ben, you also said correctly, even with human work.
The more AI is involved, the lower the yield rates.
And yield rate is one of the blockers that once you remove, you can really talk about AI leverage.
Because if you're merging 5% less of the PRs, then that is going to ding your AI leverage and the kind of production multiple that you're looking for.
That's like an immediate fine on your output.
You know, I think calling out the ownership problem is...
Really smart here.
Up until very recently, PRs were actually, shockingly, surprisingly, considering, you know, they just hold code.
They were kind of like a deeply personal thing where, like, you wrote the code and you want it in the code pace and you're going to go hunt down your reviewers.
You're going to go ping somebody on Slack or Teams and be like, please review my PR.
And then they're going to look at it and say, looks good to me.
And it's like a good, deeply, like, human.
relationship about, oh, we're shipping code, we're doing good stuff.
And you take the review personally and you take issue with the comments that you receive.
Yeah, exactly.
Yeah, exactly.
The little nits.
You don't want to ask that person because they always pick on this kind of style thing that you're not trying to pay attention to.
Right?
There is a deeply human element to the PR process.
A lot of that has changed really dramatically with introduction of AI workflows that do pick up these maybe lower priority things and cycle through them.
I think it speaks a lot to That backlog, as it so was, like all companies have this huge backlog of things that they would love to do if they had a million hours and nothing else to do and just like burn through and get it all done.
But then you actually have the ability to do it and you start throwing an agent at it and you start realizing, oh, actually, maybe.
There's a low yield rate on this stuff because maybe we don't really know what this needs to be.
And this was a placeholder.
Or maybe that this was just such a small consequential thing that it doesn't relate to the bigger problem.
And we shouldn't have even have wasted energy on it.
It actually calls attention to like what now qualifies to go into the backlog.
And for many teams, the backlog doesn't exist.
If it's not something they can immediately delegate out to an agent to run through a cycle, then it's not ready to hit.
those kinds of systems yet.
We need to be talking about them more.
And so I think that has exposed a really interesting change in how PRs work.
I also think that PRs up until recently have been very much kind of like a receipt as well, right?
It's like, this is where code is going to hit production.
This is where the, like you said, the tests are going to run here.
This is who reviewed it.
This is all of the...
the safeguards that went into protecting it.
Nowadays, like a PR, whether it's authored by a human or authored by an agent, you go on there and it's just like, in some cases, just.
tons of context is dumped on here now, sometimes way, way, way, way more than was ever even dedicated to writing the PR.
Maybe the PR is just changing a line or two or adding some stuff.
And all of a sudden you got like a whole bunch of stuff running and people on here.
And that also contributes to the cognitive load for folks to turn to and be like, oh, yes, this is ready to ship.
It makes it harder to cut through a noise.
And so a big part of that too is.
Having more eyes on the review process because we have to acknowledge that it's evolved beyond this deeply interpersonal thing and it's more agentic now.
Let's fight fire with fire.
And in some cases bring in agents to help obviously different agents that aren't the same ones that wrote the code to review it and to look at it and provide input before a human enters the scene.
Maybe to kind of give it.
a bit of a bump.
And I think this is actually something that is correlated really strongly here in the benchmarks about the bumps in this yield rate with code review.
Do you want to talk about that since you were looking at this data so closely about how even things like having AI review your code before a human can give velocity bumps and kind of help address this problem?
Yeah, definitely.
So those PRs, which used to be personal, like you said, even when I'm fully in ownership, It's no longer my code, right?
I just, you know, Claude wrote the code for me and I'm just, I'm the one pushing it.
I'm the one owning it.
I don't get the comments on the code.
It's not comments on my code.
So that dynamic is shifting even when I'm totally in charge.
But reviewing code, at least a large part of it can be done by AI, right?
It's already done by, you know, multiple solutions for AI code reviews.
It could definitely catch.
everything that's basic, everything that's about behaviors, and everything that is almost like linting, but it also can catch difficult to find bugs.
I'm not going to go into the discussion of whether developers need to look at code at all, or is that in our future where code becomes an abstraction?
But even today, when you pull a human to review an incoming change to the code base, an incoming PR, You can use AI before that human to make sure that the PR I'm looking at is now in a great state.
And you reduce the human work, which, you know, is always a huge bottleneck.
You improve my well-being because I don't have to look at messy and sloppy code.
That's already been fixed by the AI code review and the cycles AI to AI or AI to human.
And I can focus on what's important.
There's still some...
some hurdles.
For example, a heap of context in the PR because everything, typically AI, just dumps everything.
Everything in human history about is now I have to read through that.
So I think there's teams still need to find ways to improve the AI being too talkative.
And it's so easy to generate a huge design document.
So let's just generate.
If you want a human to actually read something, there is like actual hard work to get it.
terse enough, but still meaningful, that a human can actually parse it instead of just going, you know, nodding over the huge body of text.
I think if the actual code is becoming easier, right?
You can focus on small PRs and the actual code, having a human look at it is still something we can do.
Having the PRs is a meaningful kind of ownership.
transition or a meaningful interface for this is a change that we want to inject into the code base.
We can merge it.
We can roll it back.
There is good context from the tooling that ran around it that looked at this point in time in terms of the code base and gave me the verdicts.
And then I have the human at the end, the expert where needed that's saying I also bring my own expertise, my experience with the code base.
to bear when looking at the code, but I don't have to do nitpicking or look for silly bugs.
That's where the magic happens.
And like you said, having AI code review as part of my process, our data shows it, like bumps up the yield rate across almost all kinds of PRs in at least two or 3%.
It's just a higher percentage that this PR will actually get merged if the first step was an AI code review.
Before we wrap things up today, I want to talk about how we're looking at, because a lot of these discussions are directly influencing how LinearB is building our product.
We're building stuff to measure and track this and to understand which teams are being successful with AI and which have the biggest opportunities to improve.
We've talked about a lot of metrics today.
I think one thing that right now, like in the moment, it feels kind of like everything's kind of boiling down to like cost per PR.
To be like, you know, like how many tokens did you spend creating that PR?
Like that's kind of like sort of where things are gravitating towards in the moment.
We've seen this both internally at Linear B with our own engineering team, but we're also seeing it across the organization of this almost like halo effect that you can sort of tap into to get more developers operating in this agentic way.
So all the people who are listening to this right now and wondering like.
I understand why I need to be measuring this stuff and I want to be able to take the next step to improve our productivity from it.
You know, how should they be thinking about, you know, the metrics they're tracking, but more importantly, the improvement that they drive off of them?
Yeah, so you mentioned the cost per PR and I think that's, especially if I'm looking at a larger organization and like it or not, this is kind of a factory, right?
You have people and you have tools and electricity or tokens coming in.
And you have output going on.
It's always been hard to measure the output in terms of business value.
That's always been the case.
But if you're looking at your, you know, the PR merge rates, that gives you a good notion of productivity.
If you look at your whole factory, you're going to look at, okay, how many PRs is this generating over time?
You need to assume people are working on the right things.
But prioritization aside, that's my productivity.
That's my output.
If you look at the cost per PR, and here we will blend both the human cost and the AI cost.
So overall, I'm, you know, delivered a thousand PRs this month and it cost me a hundred thousand dollars.
That's a hundred dollars per PR.
And if you're driving the cost down and increasing the productivity, you're done more with less or more with the same.
And that represents the leverage you're getting from this new technology, this new tool, this new kind of electricity.
running through your factory.
You unblock hurdles like the yield rate or the, you know, the dropping quality that sometimes the AI brings.
That will drive down the cost per PR directly.
You improve your throughput by doing more in parallel.
That will drive down the cost per PR.
So it's a very natural way for managers to look at overall.
I think it's less valuable to look at a specific PR and say, how much did this PR cost me?
But...
If I'm looking at my overall, what's the team doing?
And you're saying, okay, I have two different groups in my organization and the cost per PR is dramatically different in those.
And the trend is not going in the right way.
Then I know where to focus.
I know where to focus my efforts on unblocking.
And that's the second line of metrics to help me understand the bottlenecks.
And okay, where is our problem?
Why is this cost not going down?
The cost is going down and productivity is going up.
How do I put more juice into that team?
By the way, that could be hiring more people to that team as well.
It's a good functional team that works well.
Give it more juice with people and with tokens.
And the other teams will help them solve the problems that are now creating a kind of a floor for the cost per PR and not allowing it to drop further.
Well, Yashar, it was so great having you back on the show, and I'm certain that we'll have you return again soon enough.
You know, listeners, you can get all of this data that we've been discussing today in our report, the AI productivity gap, how elite AI teams are leading the pack.
And it really gives the full story on how this productivity gap is appearing and the metrics that you need to be tracking today as an engineering leader.
to make sure that your teams are going to be driving successful AI adoption and leverage, not just adoption.
So follow Dev Interrupted.
We're on LinkedIn.
We're on Substack.
We are on YouTube.
And if you enjoyed this, you know, give us a rating or a comment or reach out to us on social media.
We always love the interactions with our audience and it helps us grow the show.
So we'll see you next week.
Thank you.
It was great.
