# DX Research: AI Engineering Gains Modest, Culture Key

**Podcast:** Engineering Enablement by DX
**Published:** 2026-06-08

## Transcript

Welcome back to the Engineering Enablement Podcast.
I'm your host, Justin Riach.
This episode was recorded live at DX Annual, where we were joined by Abhi Noda, co-founder and CEO of DX, alongside Brian Hauck, co-author of The Space Framework, and more recently, Distinguished Scientist at DX.
As part of the opening session at DX Annual, Abhi and Brian shared early findings from DX's latest longitudinal research on AI's impact on engineering throughput and developer productivity.
They discuss what organizations are actually seeing as AI adoption scales, why productivity gains are often more modest than expected, and where engineering leaders are still struggling to connect AI adoption to measurable business outcomes.
Let's get into it.
Excited to be here today.
The lineup is incredible.
There are so many interesting talks that I cannot wait to sort of dive into.
But first, we're going to dig into one of my absolute favorite topics.
Abhi and the team at DX have been working on a research report covering AI's impact on engineering velocity.
It's not published yet, so we're all going to get a little bit of a sneak peek into some of their early findings.
Now, before we dive in, I do want to say we only have about 30 minutes, so we don't actually have time for Q&A, but if there are some burning questions you have, please find Abhi or I afterwards, and we would be happy to talk about it in more depth.
All right, so to kick us off, Abhi, you are sitting on one of the coolest, largest data sets on how engineers actually work in the real world.
And this upcoming report is digging into what is happening as AI adoption really scales across our industry.
So before we get into the findings, I'm interested in what motivated this research, what really sparked it.
Yeah, thanks, Brian.
And I just want to echo Grayson's welcome.
This is so exciting to finally have everyone in a room to be able to gather and learn from one another.
So a little background on this study we've been working on.
Nearly every conversation I've had with folks like you over the past few months has begun with folks expressing that, you know, my CEO, our executives are expecting these astronomical gains.
Because there's so much hype about AI in the media and hearsay.
It's not what we're seeing on the ground.
How do we close that gap?
How do we set realistic expectations?
How do we even know what to be aiming for?
And we're hearing this from, I bet everyone in this room has probably experienced this to some extent.
And so at DX, we started asking ourselves the same question.
Like, what really are companies seeing?
Like, what can we actually see in the data?
What is the impact we're seeing?
And so that's what we've been digging into at DX.
And this is just the beginning.
As we'll talk about today, there's a lot more to unpack.
There's a lot of nuance in the data.
But at least we have sort of a first glimpse at what we're seeing at least over the past 12 months as AI adoption has really matured and soared across companies.
Awesome.
Well, I can't wait to dive into some of the findings.
I'm a researcher, and I love nerding out over research methodology.
And I don't want to spend too much time on this, but I'd like it if you could just walk us through a little bit.
How did DX actually investigate this?
How did you design the study?
And how did you decide which companies to include in it?
So at DX, of course, our customer network is, at this point, about 500 companies.
For the purpose of this study, we took a sample of those companies just because of the...
The effort required, we couldn't really include every company.
What we looked for in terms of the selection of the companies were companies that have reached a point of maturity in their developers' adoption of AI.
So we defined that as, I believe it was over 75% monthly active usage of tools.
And in almost all cases, there's been a significant rise in adoption over the past 12 months.
We also looked and narrowed the scope to companies with over 100 engineers.
We also excluded companies that had gone through major liquidity events or M&A, IPO, major changes where regulatory impact could have a play.
And in terms of the methodology, it's really hard.
It's really hard to single out the causality in terms of what is AI and what are the other confounding factors.
One of the things I see...
folks do a lot is cross-sectional analysis of, hey, the developers using AI are more productive than the ones who aren't.
And there's flaws, I think, in that approach, because oftentimes the developers using AI were the ones who were already coding the most.
And so, you know, they always look better.
In our case, what we did is look at a, you know, longitudinal data.
So we looked at the organizational level as AI adoption has matured.
What has been the impact on the organizational velocity and throughput?
And we can talk more about what we mean by throughput and velocity in a moment.
Yeah, I love that.
And as somebody who has been studying the impact of AI on the industry, being able to control for all of those different biases are so challenging.
And so I'm glad that you did a very thoughtful approach.
So I'm curious, as you set out to do this research, what were you expecting to see?
What were some of your hypotheses?
did you actually see?
Were those hypotheses correct?
So as a developer myself, I had some sort of like gut feel hypotheses, but to be totally honest, from a study standpoint, we really didn't know.
We were really eager to just see, like, I really didn't know what the answer was, and we were really pushing the team to, you know, what's the answer, what's the answer, we want to know, we want to know, just as much as you guys probably want to know.
And, you know, what we ultimately found was, you know, the TLDR is that, you know, most organizations fall kind of between that 10 to 15% mark in terms of, and we'll talk more about why we chose PR throughput as the indicator of velocity in a moment, but about 10 to 15% increase in throughput.
And the actual median was around 8%.
The mean was around 11%.
We can talk more about that in a moment as well.
So it was much more modest than I think we expected.
I think it's much more modest than many CEOs who are talking to their CEO friends about the 10x gains that the companies are delivering would expect.
I think in terms of our engagements with different companies and customers, it wasn't completely surprising.
I mean, we'll talk about why the gains are more modest than what all of us might have expected in a moment, but there's a lot of factors underlying that.
All right.
Well, to kick us off with maybe a spicier question this morning.
So one of the central themes of the space framework is that developer productivity is nuanced, it's complex.
It's about so much more than just the count of activities.
And so given that, why did you choose PR throughput as sort of like your measurement?
And do you have anything you could say about sort of the nuance of how you measured that?
Yeah.
So, I mean, of course, measuring velocity or productivity is in of itself really challenging.
And so for this study, we actually did look across a number of different metrics, more across the full spectrum of space or the DX Core 4.
But in terms of the focus of what we wanted to sort of focus on and publish, we felt like PR throughput was the most relevant right now and practical for the world, mostly because that's what we see most organizations talking about and focusing on.
So we kind of just wanted to meet the world where it's at in terms of how it's thinking about this problem.
We also explored our proprietary metric true throughput.
which is adjust PR throughput with AI to sort of like weighted PR throughput, so it takes some of the noise out.
Again, we liked that signal, but for relevance and practicality, we felt like PR throughput is a more accessible metric to focus on.
We have more data on the other metrics and how those have been affected as well, so we'll be eventually publishing those as well, but really it's about practicality and relevance in the moment, why we focus on PR throughput.
Awesome.
So, you know, I suspect for many of you in the audience, you know, your lived experiences, you know, 5%, 7%, 10% increase in PR throughput probably matches, you know, what you were feeling.
That probably rings pretty true for many of you.
I see some head nods.
But as you mentioned earlier, like when I go and talk to business leaders, they're often expecting 20, 30, 40, 50% increase in sort of quote unquote productivity.
And so I am just curious a little bit, as you went and talked to developers, like, why do you think we are seeing those gains might be lower than, you know, some people, particularly those on the outside of the industry might expect?
Yeah, so our team went and conducted follow-up interviews to really ask the question, hey, why aren't the gains higher than what you're seeing?
And there were a number of different...
These are sort of the coded categories of responses.
And this isn't the full list, by the way.
These are just the top five.
You know, the top one probably wouldn't surprise most of us in this room.
It was that coding is not the primary bottleneck for engineers.
And Brian, of course, at Microsoft, you guys pretty recently published a study on where engineering time is spent.
And I believe it was 14% of developer time is actually spent coding.
very small part of where our time and money is going.
And so if AI is only optimizing that, it partially explains maybe why the enterprise velocity gains are limited.
We also heard about how these AI tools and new automation introduces new bottlenecks.
And we all are thinking about this, things like review and technical debt, cognitive debt, new concept.
We also heard about challenges with adoption.
So it was pretty interesting, like a lot of social friction, you know, like kind of cultural clashes around adoption that's sort of inhibiting full adoption.
And there were a few more, just for brevity, I won't go into all of them, but those were some of the major themes we heard about.
So I did have a chance to preview a little bit of the data.
And, you know, we see that, or you saw rather that sort of the typical gains in throughput were 7.7%.
But you don't have to go that far out in the distribution before, you know, there were some organizations seeing substantially higher gains.
And, you know, I'm always interested in where do we have some wild outliers?
And given that, I'm curious, as you look at some of these outliers, like what set apart some of those companies that might have had some disproportionately high gains from the more typical experience?
So we don't have a great answer to this yet.
You know, we really focus on why aren't the gains higher was the first question.
The second question is, okay, for those for whom the gains are higher, why?
And I think there's sort of two parts to that question in terms of how we want to approach it.
One is that we're planning to first tease out for the companies whose gains are measurably higher.
We just want to double click into that to make sure, you know, are these superficial gains?
Are these real gains, right?
We want to peel back the onion a little bit.
And then two, we want to understand, okay, assuming these are real true gains, what are the strategies they're employing?
And by the way, we have speakers today, I think we'll be sharing some stories.
There are good success stories and strategies that will be shared.
I think a lot of it boils down to what I've seen as sort of an all-in culture, a fully bought-in culture around centralized rollout and championing.
of these AI tools.
Obviously something you work on at Microsoft with your counterparts, but that's the biggest thing I've seen.
It's more of a cultural shift toward really infusing AI into the entire AIS DLC, not just coding, and really aligned sort of cultural push to push that adoption through.
Yeah, that actually aligns with some recent research that I had published that is, you know, that showed that even things just like leadership advocacy really increased the gains that people saw from AI.
So, you know, you said something earlier that I think is really interesting and is often a misunderstood point about software engineering.
Like software engineering is...
so much more than coding.
Coding may be the things that us as devs, we want to do.
It is a central part of our identity.
But to your point, it's only about 14% of our day.
And one of the things that I've been seeing in some of my research is, well, yes, we are producing 5%, 7%, 10%, 15% more pull requests.
We are doing it substantially more efficiently.
So we are actually seeing for hands-on keyboard, time spent coding.
we're actually producing about 40% more PRs per hour of coding.
And so that hints that we are reclaiming some of that coding time.
Do you have any idea where that's getting reinvested?
Some of it clearly into more throughput, but where else is that maybe getting spread out in the SDLC?
So this has been a similarly perplexing question.
So for the past six months, I've heard customers asking, look, we're seeing X percentage self-reported time savings from developers.
They're saying they're saving this much time, but we're not seeing that show up in our output and throughput.
So what's the disconnect?
Where is that time going?
We don't know.
This requires additional research.
But I will say, you know, I think one preliminary hypothesis is if we anchor to that 14% number, if developers only spend 14% of their time coding, then it would make sense that the time they're recouping, is ratably distributed across their activity.
So only 14% of the time savings go back into more code, right?
Which would explain why we don't see, you know, one-to-one output gains mapping to time savings.
I think another hypothesis, if we, you know, refer back to the question of why aren't the gains higher, is that some of the time savings actually come with side effects that require more time.
So more time, overseeing the work, QC, QAing the work, reviewing the work.
I think there's also some anecdotal, like we've seen in our research a little bit, that there's also the lag time.
So like developers using these tools, like what do they do while they're waiting for the agents to produce the code?
Do they go play video games, go for a walk, right?
So yeah, those are some of the hypotheses, but we don't have a clear answer yet.
So you actually just started touching on it, but I am curious, like, there is a reason that we don't just measure lines of code as a good measure of productivity.
Like, just writing more code isn't always the right answer.
And I'm curious, like, as we are increasing velocity, what are some of the potential unwanted side effects that you've been seeing?
Quality and cost are the two that I hear our customers talking about constantly, right?
Cost is...
Something we're all trying to wrangle right now or just stay on top of, get some visibility into.
And quality is something that I think we all understand is sort of an underlying risk.
But I think it's still early days.
We did see AWS and Amazon go very public talking about this.
But there's some delay between the technical debt that we're perhaps creating and the later consequences that may follow.
I think the biggest thing that I've been talking to customers about is sort of this cultural risk of, I've been calling it false velocity, right?
I see this in so many organizations right now.
And to some extent, even us as developer productivity leaders can fall into this trap of just being so focused right now on showing off how much faster and prolific we are.
with AI and not really focusing on what meaningful improvement is this leading to.
So what I mean by that is, you know, you probably have engineers in your organization showing off really crazy things that they can do with Claude or whatever tool.
But we're not asking, okay, is your team's product velocity actually increasing?
Right?
And leaders are talking about, oh, how many more PRs and how many more lines of code we're generating, but is their roadmap actually accelerating?
And is the quality of their products in code?
actually sustainable.
So I think there's this risk right now of just being so focused on showing off what we can do that we're not paying attention to like, are we actually getting better?
Are we actually materially improving our businesses?
So that's like really interesting to me because one of the things that I've been trying to look at is like, what is the actual innovation velocity as a result of AI?
Are we just shifting bottlenecks around as you previously hinted at?
Or are we actually able to deliver innovation faster?
And so What's sort of your recommendation to engineering leaders who want to look at where they can put AI in lots of places in the SDLC, not just code execution?
I mean, that's the biggest thing when I talk to leaders and hear about where their heads are at right now.
A lot of folks are at a point where you've rolled out one or more of the popular coding tools to developers.
Adoption is at a fairly satisfactory place.
Maybe you're kind of double-clicking into that a bit more.
So a good amount of folks are asking, what next?
And especially if you're sitting at that kind of 10% number, you're asking, okay, well, we wanted 10x.
How do we get there, right?
I think the things I'm hearing about and seeing are, you know, one, very clearly looking left and right of code.
So we've solved the coding part to some extent.
We've accelerated that.
But as we know, if that's only 14% of where our time and money is going.
What about the rest of the 86%?
How do we accelerate and optimize that?
So I think that's a really important question.
In addition to thinking about the SDLC in terms of like left and right of code, there's also how do we improve the developer experience more deeply?
So especially if you're a DX customer, right, you're thinking about things like deep work and documentation, these friction points that developers experience.
How can we leverage AI to materially improve those things?
And the third theme, that I'm seeing a lot of organizations focus on is this idea of, there's so many names for it, autonomous engineering, async engineering, background agents, you know, this idea of, I think today we're still at a point where most of us are focused on leveraging AI to accelerate the human work, right?
Like humans are still in the cockpit, they're steering, they're monitoring the work of the agents.
How do we complement that acceleration with...
augmentation which is the idea of more autonomous agents that are truly augmenting your human workforce and working in parallel not underneath our human developers so those are kind of the three big themes and of course using data to help kind of inform and guide these investments is really important So you said a word there that I think is incredibly important for a lot of my research, human.
For those of you who might be familiar with my work is I really enjoy looking at sort of framing the human context of how we work.
And so I will look at things like how does having plants in your office make you more productive?
How does access to sunlight change your productivity?
Even things like how does, you know, exactly, like how does spraying lavender in your face while you sleep improve your productivity?
You know, you and I are both huge fans of Dr.
Margaret Ann's story, and she's been talking a lot about cognitive debt and sort of the human cost of some of these AI transformations.
And I'm curious, you know, what have you looked at in sort of that area?
I'm a big fan of that research and those ideas.
Margaret Ann's story, Peggy, has recently published.
I think that falls under what I was talking about earlier in terms of...
the risks and how I think there is also a time delay.
Like this is all, this is also recent and new.
Like we're not thinking about, okay, what are the consequences 12 months from now?
And if we adopt certain ways of working.
So absolutely, I think the loss of sort of human understanding of the systems that we're building is a really interesting risk.
There's sort of different takes on how material of a risk that is, depending on how you look at it.
Some people might argue, well, look, it doesn't matter how, if the humans don't understand the systems because they can just use AI to quickly, you know, regain that understanding when needed.
That's one argument.
You know, others may argue, no, like, if you need to call the mechanic, they better be able to, like, work on the machine and have mastery over it.
So I don't think we know yet, but I think it's a really valid and interesting idea.
So we've sort of hinted at it throughout this talk, this notion of are we actually delivering innovation faster or are we just moving bottlenecks around?
And sort of my first reaction for any time we have an unanswered question is, well, how do we try to measure it?
And I'm curious if you have any thoughts on how should we actually be measuring if we're just sort of shifting our bottlenecks?
Yeah.
And as everyone in this room is probably thinking about and as we at DX have been talking about, I think our approach to measurement, some things say the same and some things need to evolve and are evolving, right?
You know, our perspective has been that this is a framework we published, I think it was last September.
So it feels really old in my mind.
But, you know, I think that much of this still holds true today.
There's also a lot that has changed and that is new that we've been working on at DX.
But, you know, to some extent, how we think about the overall software organization and things like quality and velocity, I think, you know, stay the same.
And especially if we're trying to understand how are things changing with AI, you need consistent measures pre post during this transformation.
But there's also a lot of new tools, new ways of working, new workflows, many new workflows being born every week.
And so how do we sort of measure these new ways of working?
Well, those do require new approaches.
And I'll touch on two things, and we can maybe double-click into them some more if we have time.
You know, one is, I think, increasingly, I think there's a need to separate out how you're measuring.
So you're really thinking about that acceleration.
So the human lift, like how much faster are our humans is one bucket.
And then the augmentation, like how much more capacity are we generating and creating with agents?
I think thinking of those as two separate.
buckets in your overall formula is really useful.
And then the question of how do we measure agents is really interesting.
How do we measure ROI?
We have some ideas here that we've seen out in the real world work pretty well.
So this idea of reducing our understanding of agents to agent hourly rate, which kind of gives us this return on investment idea that we can compare to, say, human hourly rate.
which gives us an interesting comparison point.
Something really new that was not even in our minds last September, but some of you may have heard about, is this idea of agent experience.
So in the same way that DX was founded, this whole idea, how do we measure developer effectiveness?
Well, you go to developers and you get feedback and signal from them.
At DX, we've just released a way, very similar idea, but for agents.
How do you measure AI?
agent engineering effectiveness, well, you go to the agents and you get feedback from them on where are their bottlenecks and constraints.
So we've just rolled that out.
We're surveying agents.
That's pretty crazy.
Okay.
Like that is wild today.
And I know we're almost at time.
So like, do you have like one bullet, like what's one finding you have on how to make agents more effective?
Well, you got to ask them, just like we ask our developers.
So, you know, it's really early days, but.
We've initially begun with measuring four factors, and I won't list all of them, but for example, we'll ask the agents, how were the requirements that you were given?
How easily were you able to understand the code base that you were working within?
How well were you steered by the human that you were pairing with?
So these are the types of questions we're asking agents in a calibrated way, so we get quantitative.
metrics from that, as well as qualitative feedback.
So the agent will explain where they had challenges or bottlenecks.
So that's the rough idea.
I love it.
I love the same principles to measurement apply.
So that is, unfortunately, all the time we have.
We were able to cover a lot of ground, thankfully.
Abhi, thank you so much for sharing some of these early sneak peeks into your research.
Can't wait to read the full report.
You had any questions that we weren't able to get to today, please find us throughout the day, particularly Abhi.
He's the man with all the answers.
And just thank you so much.
And with that, I will pass the mic back to Justin, I think, to introduce the next session.
Thank you, Brian.
Thank you so much.
