# Optimizing Agent Experience and AI Readiness

**Podcast:** Engineering Enablement by DX
**Published:** 2026-07-24

## Transcript

I think that context quality is where the big lever lives right now.
Is the context that agents are working for unclear?
Are the tasks actionable?
Is the data accurate?
You can't do it simply by trying to improve the agent itself.
You improve it by investing in the conditions around the agent, just like with developer experience.
Welcome to the Engineering Enablement Podcast.
I'm your host this week, Justin Riach.
Before we jump into the episode, I want to tell you about the AI Measurement Framework.
a research-based set of metrics designed to help you understand the impact AI is having on engineering productivity.
To learn more about this framework and how to apply it, head over to getdx.com slash AI dash framework.
Welcome to another episode of the Engineering Enablement Podcast.
I'm Justin Riach.
While engineering leaders navigate rapid changes in developer workflow as a result of AI adoption, understanding how these changes impact the developer experience is critical for maintaining performance, quality, and culture.
In this episode, I'll speak with Brian Houck, Distinguished Scientist at DX and Prolific Developer Experience Researcher, about his current research into the intersection of developer productivity, agent experience, and AI fluency.
Let's get into it.
Thanks, everybody, for joining us today.
I am really excited about this conversation.
You know, Brian and I have been in this space, you know, looking at developer experience and productivity for a number of years.
In fact, I was just pulling up an old...
podcasts that we did together almost five years ago now and it was amazing to me like that just the difference in the context of the things that we were discussing back then you know looking at tying developer joy to developer productivity the first principles didn't shift very much you know some of the metrics that we thought were important and certainly the the spirit of developer experience but you know who who would have thought that almost five years later here we are we're gonna have a conversation we're gonna delve into agent experience today so Anyway, thanks for joining.
We're going to go through some of the newer research that we've seen with DX.
I'll be kind of playing host, and Brian will kind of be like panelists.
We've got some questions that we're going to go through, but we both have some opinions on this, so we'll be sharing our perspectives and insights.
I'm Justin React, deputy CTO at DX.
Nice to see you all.
And Brian, why don't you go ahead and introduce yourself, and then we'll get things kicked off.
Awesome.
Thanks, Justin.
Yeah, so I'm Brian Houck.
I'm an applied scientist on the DX research team.
I'm fairly new to the team.
I take a very broad view to the study of developer productivity.
So I do a lot of research exploring everything from how tooling factors like AI are impacting developer experience to how collaboration patterns change developer experience, even to like wild things like studying how like sunlight impacts developer productivity.
And a lot of my work centers on measurement frameworks and metric design.
And so I'm just super excited to be here, geeking out about developer productivity with you.
Oh, we're going to have a blast.
I always learn so much from you, Brian.
And yeah, you talk about things like sunlight.
I mean, so many variables.
It's what makes this such a nuanced.
problem right let's kind of dive right into it we will be watching uh the qa and chat for questions and we do have somebody helping to moderate those so as much as possible i'll do my best to move to audience questions as they come in so feel feel free to interact with us we we like that makes us remember that we're not just screaming into our laptops here so uh let's kick it off talking about asian experience we spent a lot of time talking about developer experience that's what we measure that's that's the acronym for our company but what In your mind, Brian, is agent experience as a research area specifically?
And why should engineering leaders care about it, not just researchers like us?
I think this is such an interesting emerging area.
So I think of agent experience as a research area that explores whether agents have the optimal environment to do their best work.
Do they have clear requirements?
Are they working with accurate data?
Do they know what good output even looks like?
I think it's the natural extension of developer experience in a world where agents are becoming first-class sort of participants in the engineering system themselves.
And so if we want to maximize sort of the human-agent collaboration...
Developer experience tackles that from the developer side, and agent experience tackles it from the agent side.
And engineering leaders should care because if agents aren't set up for success, then humans can't produce their best work.
And I think part of the challenge is many leaders evaluate AI tools on tool capability.
Does it write good code?
Agent experience, I think, helps answer the harder question, like, does the human-agent collaboration produce good outcomes?
And I think, you know, one thing that has stuck with me is I recently published some findings from a great researcher, Sarah Chisari, in a paper, The Space of AI.
And Sarah found that the strongest predictor of multi-agent success was not tool capability.
but specification quality, like how well humans set context.
And so, you know, the agent capabilities is what the vendors sell you.
The agent experience is what the developers actually get.
And those aren't the same things.
And that's sort of what I've been looking at there.
That's like super interesting because there's so many facets to that too.
I mean, I think that like when we...
we talk about like you know what's good for humans is also good for agents and so we can certainly to a degree use what we've considered to be good developer experience in the past as a foundation for building good agent experience but there is another dimension right i mean there is like you know how skilled is the engineer in steering the agent how well have we provided context how prepared is the platform to provide that context so What are some things that you're seeing in terms of efforts to really improve the agent experience?
And what have we seen so far in terms of the improved outcomes that you mentioned from that?
Honestly, I think we need a lot more research in this area.
But to figure it out, we should borrow from the playbook that worked for improving developer experience for years.
It's define it, measure it, improve it.
And so if we try to break that down is, can we come up with shared definitions, shared principles for what a great agent experience even means.
Like that's the starting point.
And then how do we measure it?
It's much like a great first step in measuring developer experience is to ask developers.
I think the first step of measuring agent experience is to ask agents.
And trying to figure out how to effectively survey agents about their agent experience is like something I'm spending a lot of time trying to wrap my mind around right now, in fact.
But ultimately, like that is to give us the signal for how to improve it.
And I think that context quality is where the big lever lives right now.
And so is the context that agents are working from clear?
Are the tasks actionable?
Is the data accurate?
And I think things like stored skills and shared prompts help, data validation is critical, but also clarity of intent.
Like, have you really thought about what you're actually trying to accomplish?
And I do think it's important as we try to improve agent experience, you can't do it simply by trying to improve the agent itself.
You improve it by investing in the conditions around the agent, just like with developer experience.
That's interesting too, because certainly the models are continuing to improve and we're able to store more parameters and we're able to shuttle more data into these things.
But the big...
you know, epic shifts in the capabilities of these models have not really been, you know, in like improving the underlying LLM, but more like how are we providing context, you know, things like RAG and MCP and I mean, all these architectural adjustments.
And so I think that's right, right?
I mean, when we really see like leaps and improvements here, it's more about, you know, how can we efficiently provide like this additional context kind of to the agent?
And what should people be thinking about?
I mean, we're...
I think, you know, those who are very familiar with like improvement in developer experience are probably, you know, kind of clued in into what some of these dimensions of improvement are.
But can you give me an example of a few of these things that we should really be focusing on at the platform level when it comes to improving agent experience?
I touched on some of these things on.
You know, do you have a library of shared prompts and things like that?
It's as I think of some of the dimensions, you know, I've already touched on them, like clarity is the context you are giving to agents.
Is it.
unambiguous or is it up for interpretation?
Because if it's up for interpretation, you're going to get a lot of randomness in the results.
Is like I said, like, is the data accurate?
And so you should have like at the platform level, like, how do you do data validation?
You know, you mentioned things like RAG and I think sort of the efficiency of our human agent collaboration on things like, are we giving it the right amount of context, the right scope, or, you know, something that I am as guilty of as anyone is like, oh, I'm really asking about this chunk of code, but I pass in my entire code base because why not?
Like those sorts of things where you can put protections and guardrails in place at the platform level, I think will be super helpful.
I'm actually curious, like, Justin, like if you have heard from sort of customers across the industry, what others might be doing.
I mean, these are things like concise and up-to-date documentation.
You know, that's a big one.
Thinking about areas of delivery like fast CI, you know, that's not so much necessarily something that the agent is going to be, you know, quote unquote, aware of as much as it is, you know, improving the efficiency of agent.
If we can generate code instantly, but we have a 45 minute build time.
Well, who cares that we can generate code instantly like that's the bottleneck is going to be the is going to be the build system, right?
If we don't have if we have flaky tests and those flaky tests are difficult for validating the code that we're sticking in for human engineers, guess what?
The same thing is going to extend the agent, except now we're generating 10 times as many PRs with twice the PR size that we're seeing in some of our more recent data.
And so it's just going to exacerbate these existing problems.
The importance of having good documentation.
If only we hadn't known this for years and years as it relates to humans, and it's just compounded now as it relates to agents.
That's the irony, isn't it?
I mean, for me, and I'm sure for you too, we've been clawing at this space for years.
We've been really trying to convince executives to invest in a better developer experience.
And now we're finally doing it because of all this spending that we're doing on AI.
I've come to grips with it.
You know, it's like, okay, I don't really care what the catalyst is as long as we're finally going to take care of these aspects of the developer experience that are now extending to agents.
And I mean, I think we've already kind of hit on then, you know, this relationship between developer experience and agent experience.
But do you think, I mean, you know, maybe more broadly, I mean, this is, this does become, even if it's accidental, kind of a virtuous thing, right?
Because if we're going to be improving this for agents and we're making developer experience better, what are your, what are your thoughts on that?
Oh my gosh, absolutely.
Like, I think they are, deeply intertwined and increasingly inseparable.
Like the agent is now part of the system that the developer works and operates in.
And that by definition sort of means it's a critical part of the overall developer experience.
But I do think that there are some signals in research that I've read recently that shows that that relationship has some complicated implications.
And so a...
a phenomenal researcher Annie Vela recently released this like really cool longitudinal study that found what they called the productivity experience paradox.
And what she had found is that developers as they are increasingly likely to report that AI is making them more productive.
they are also increasingly likely to say that their developer experience is getting worse.
And for those of us, like you and I, who have spent years working with things like frameworks like Space and DevEx, it's pretty mind-blowing to say that for some of these AI-assisted workflows, productivity and developer experience are decoupling, and they were always so tightly coupled.
And so the ramification there is as we are optimizing for productivity, maybe we're getting short-term wins, you know, we're seeing, you know, PR throughput spike and things of that nature.
We might be degrading other aspects of the developer experience, and that could have...
ill tidings for longer term ramifications like retention.
It could have negative impact on code quality.
It might not set you up for sustainable productivity long term.
And so I think you can't separate the two and you definitely, definitely, definitely can't measure only one of them.
And then I think on the flip side is like as you improve agent experience, that's now a developer experience intervention as well.
Now, it's really interesting.
And obviously, without naming any names, but some of the more AI forward companies that we even see in our platform have some of the lowest developer experience index scores.
And so I think maybe you're right.
We kind of.
push, sometimes we push past some of these more human dimensions of platform improvement.
I think that's really important.
I'm glad that you brought that up because you're right, we can't disentangle them.
By the way, I'm seeing some questions come in.
We are recording those and we will be making some time at the end here to go through each of those questions.
Thank you so much for submitting those.
Okay, so maybe finally then something a little bit more actionable.
And what can...
companies do to ensure that they are ready, I mean, specifically for deploying agents.
When we kind of shift from this agent experience, which I think is the outcome of providing better readiness, AI readiness for our platforms, beyond the things that we've discussed, like documentation and better CI, you know, what can we be focused on and how really should we gauge that companies are ready for all of this agentic work?
So I certainly have a perspective.
I'm honestly going to be more interested in your perspective.
I think you might have a better pulse on this than I do.
But I think of AIA readiness as sort of hitting on three different layers.
And I think a challenge in this conversation is that most organizations only focus on one, tooling.
But I think sort of the two deeper layers are culture and infrastructure.
And the research is actually fairly unambiguous on what matters most.
And so if I sort of deconstruct that, tooling readiness, it's things like, do you have the licenses?
And do your employees have access?
Is it integrated?
Are the tools integrated into your existing workflows?
Completely necessary, but it's the easiest part.
And I think that oftentimes organizations over-index on it.
Cultural readiness is, I think, by far the highest leverage layer.
Again, the space of AI study, one of the things we found there is that developers in organizations where leadership actively and strongly advocates for the usage of AI found that daily adoption was 7x what it is for developers who are in organizations that aren't as strongly advocating for the uses.
7X is this massive change.
And I think that cultural readiness, it applies to also things like, do organizations have robust commitments to training efforts, things like peer mentoring?
A thing that we published in, again, the Space of AI study was this notion of social proof.
When a developer watches a trusted colleague solve a real-world problem with AI, that was the single most consistently cited adoption catalyst.
Training programs are incredibly important, but it is a distant second to this, like, can I see what my peers around me are actually doing this social proof?
And that's all part of sort of cultural readiness.
And then the last thing is infrastructure readiness.
And this comes down to measurement.
Like, before rolling out new AI tools or workflows, do you have a baseline?
of where your system is today in order to compare against?
Do you have feedback loops with your developers so you can hear how it's going?
And I think most organizations or these many organizations skip this and then you can't tell if anything is working.
And I guess the last thing, I said there are three layers, but I'm going to have like a bonus layer here for the agent side of this, for agents specifically, which is context readiness, right?
Like agents amplify whatever sort of specification quality already exists.
And if a team operates without clear specs, agents are going to produce sort of a nightmare of results.
And I think like those three plus one dimensions are what ladder up to readiness.
But I'd love to hear like what you've seen and heard.
I think that you gave a near perfect answer there, honestly.
I mean, I think, well done.
We can go home.
We're done.
We're just great.
So no, apart from the platform readiness, which, you know, are these things that we've already mentioned, good DevOps practices, things that we've learned first principles over the last decade or so.
There are certain, like, you know, like agent-specific things to, like having agent markdown.
present separate from the human readable documentation, making sure that like we've really got our linting and our style checks and things like that.
You mentioned setting up these feedback loops.
I think that's really important.
Like one of the most important feedback loops is having somebody shepherding and maintaining the durable prompts that are being sent, the system prompts that are being sent alongside other prompts that engineers are putting in.
And it's really like, I don't care how you do that.
You know, it's the feedback loop is what's most important.
So whether you've got it.
a ticketing system or a Slack channel or even opening up some of those durable prompts to source control so that engineers can make suggestions to what the durable prompt should be looking like.
It's more about having somebody gatekeeping that, making sure that there's an open channel for when the agents are doing something that they're not supposed to do, that we then put that as part of the system prompts so that the agent can correct its behavior kind of going forward.
And I love that you mentioned cultural readiness, specifically around areas like psychological safety, which we know, you know, you look at Google's Project Aristotle from mid-2010s.
where they found like they had a hypothesis that high performing teams would be mostly comprised of experienced engineers and really good leaders and unlimited access to compute power.
And they were like totally wrong.
Like the number one factor was psychological safety, like overwhelmingly.
And so being proactive about alleviating fears of displacement, making sure that we're tied at tying, like adding like leaders are thinking that we should be tying these skill sets to employee success because these skills are going to benefit.
employees for the remainder of their careers, right?
So I think, yeah, it's a, I would agree with everything that you said in terms of these three layers plus the bonus layer.
And I think that we can't ignore the cultural aspect and the psychological safety aspect of this.
There was a recent straw poll that I saw executed where engineers were asked to sum up in three words how they felt about you know, AI in the workplace.
And it was like very consistent to have engineers putting both like excited and terrified, like in the same breath, you know?
So it's like, let's see if we can reduce the terror a little bit and push the excitement.
I mean, like I think, and in fact, I've seen on the call of like, cognitive psychologists are going to have so much interesting work for a long time around.
Like what is just like the psychological, ramifications of all of these new tools and experiences.
Oh, totally.
Yeah.
I mean, how is it going to change our cognition?
Because our brain science hasn't changed.
How is it going to change the way our brains work, potentially?
Like much like, you know, social media has rewired our brains in certain ways, is AI going to rewire our brains?
Oh, and we spend so much time, we've spent so much time in this space warning against context switching.
And we're really trying to encourage, you know, building platforms that encourage flow state.
Now it's a feature.
10 CLIs open, you know, with 10 different agents.
And there's a term for it already.
AI brain fry is what we're calling it.
And now I think the orchestration layer is going to become even more important.
You know, if we look at projects like Gastown, Gas City, which are not ready to be used, but I think you squint a little bit and you get a glimpse of where this is going to go in terms of orchestration.
So, well, and then talking about some of those first principles, then let's talk about the core four.
Let's talk about the DX core four framework a little bit.
So obviously that it.
dates by several years, this current hype cycle that we're in right now.
But I would love to hear, you know, you were a co-author of the space framework, you know, you spent so much time, you know, researching good measurement frameworks.
What still holds in the core for, in your opinion today?
What should we be reinterpreting?
And is there anything as part of the core for that's just like broken now?
This is, this is, uh, a question I am particularly passionate about, and I feel like I'm getting it an increasing amount now from everyone I know.
And I feel like as an industry, we are afflicted with this collective impulse to throw away all of the tried and true developer experience metrics that have been helping us for years.
And it just makes me want to pull my hair out.
I am a firm believer that almost everything in the core four holds.
by design, because it's a framework designed around measuring outcomes.
And so the core four, it has four top level dimensions, speed, effectiveness, quality, and business impact.
And those outcome dimensions describe what engineering teams, what engineering organizations are trying to accomplish.
AI doesn't change what we're trying to accomplish.
It just changes the way we get there.
Like AI is a means to an end.
AI is an incredibly powerful tool set of tools, but we're not using AI for the sake of using AI.
We're using it to accomplish something.
And frameworks, durable frameworks are those that have an eye towards the outcome.
And I think what really changes is not those top level dimensions, but what are sort of the diagnostic metrics underneath.
underneath those dimensions, sort of powering them.
And so when I think about like what holds, the four dimensions of the core four still hold.
I think those are still the right dimensions.
They each have a key outcome metric that sort of ladders up to them.
So these are things like DXi or change failure rate.
And I think that protects sort of this like intellectual lineage tying, you know, DORA, which really is about delivery outcomes to space, which is really just sort of a set of principles on like, hey, like.
measure across lots of dimensions, you know, to DevX, which is really around like lived experiences.
As I think that holds, I will say there's a little bit of an asterisk around PR throughput.
I think like that still holds.
I think that's still valuable, but like that might be changing a bit, which we can certainly get into.
Now, I do think what requires some reinterpretation.
are our diagnostic metrics, those metrics that help explain what we are seeing in our higher level, top level outcome metrics.
These are things like PR merge rate, time to 10th PR.
Like the definitions haven't changed, but sort of the causal relationships around them have changed.
And so like an example is, let's say we have a 95% merge rate.
Well, does that mean that AI is producing flawless code?
Does it mean that humans are just rubber stamping everything?
Does it mean that we're not like pushing agents and finding like where the bounds of their sort of like capacities, capabilities are?
You know, we're playing it too safe.
Like it's the same number, but it can have very different implications on system health now than it might have otherwise before.
I guess in closing, I would say my very strong view, and I'm going to release an entire article about this, is that the framework is stable.
It's just the interpretation.
is where the real work begins and understanding these new emerging patterns.
I mean, I completely agree too.
I think that, you know, these foundational metrics, the way that we've been kind of positioning it, like the new metrics that we can look at, the telemetry metrics coming from the APIs, showing us like cohorts of users and things like these are useful for segregating cohorts that we can then turn around and cross reference to our foundational productivity metrics.
But they're really only like telling us what's happening with the tech.
Not like, are we actually shipping more value?
Are we actually getting more features out to market faster?
Which is really what we want to know, right?
And that's where the foundational metrics, I think, then tell us whether or not these investments are actually working.
So I know- They're input metrics, right?
Like they describe how we get the work done.
And then we have sort of engineering system metrics that describe how that work flows through the system.
And then we have outcome metrics.
And I think it's easy to get lost in the hype and just forget that this is still about developer experience and developer productivity.
And it's interesting too, when we look at some of the data that we source, like in our Q1 AI impact report and our upcoming Q2 report, which is coming out in just a couple of weeks, like you look at these patterns, like we've looked at per company impact equality metrics.
It's very volatile.
We see some people going up in quality, others going down in quality.
If we peel back the utilization from that.
you see basically the same pattern.
You just don't see the amplitude that we see now.
So we're like, you know, an order of magnitude higher now, whereas we might've shifted two points.
Now we're shifting 20 points, but the underlying pattern is still representative of the health of the platform, the health of the engineering team, you know, all the foundational metrics that we've been able to trust up in history.
Like all of the core principles of what good sort of development patterns have always been is just like AI supercharging.
And that's such a cool finding.
It really is interesting.
I mean, when you look at the sway between companies, it should be a little bit alarming, just kind of a reminder that we really do need to measure and that we need to just stay vigilant as we kind of turn up the gas on all these things.
And I think that we will build some new metrics too.
I mean, like we were already talking about agent experience.
I think that there are metrics to be found in there, you know, the effectiveness and quality of our prompting and our context engineering.
And these ultimately will be tied to performance.
I think we've been throwing around the statement that AI is not going to take your job, but somebody really good at AI will probably take your job.
And so it's good to be good at AI.
And so some performance metrics there, I think, will be helpful as well.
Kind of along that same thread, we've had people reach out and ask questions about whether PR throughput even makes sense anymore when agents are acting as sort of extra humans on the team.
What do you think?
we should be, how should we be interpreting PR throughput now?
Is this one of those metrics that we should be reinterpreting?
What are your thoughts there?
So like, I will acknowledge PR throughput has always been a controversial metric.
And I will say like, I am a strong defender of PR throughput historically as a great measure of how quickly and easily can we flow work through the system?
I think it's a garbage measure of sort of individual performance.
There's too much normalization that has to happen.
But if we zoom out as a system level metric, I think he's always told us a lot about the friction in our developer experience.
Now, that being said, I do think that PR throughput sort of occupies a unique position amongst sort of the key core four metrics because it is directly tied to the mechanics of software delivery.
And that's like what AI is changing the most.
And as more and more code is being written by AI, which is something like...
you've been looking at a lot lately, it's going to be an important measure of can we still flow that code through the system, but it's not going to tell us very much about developer work.
And I think particularly in the future where sort of the atomic unit of work starts shifting away from being PR-centric to being what some people are calling intent-centric or specification-centric, the unit of work changes and so does this metrics meaning.
I think long-term, like we're not there yet.
And so I think PR throughput is still both valuable as sort of like diagnostic, but I think it's still valuable as an outcome.
Like, are we flowing work through our system successfully?
But I think longer term, this sort of gets moved to a purely diagnostic supporting metric and sort of our top level speed metric looks at the flow of innovation more broadly.
And so I recently published a paper.
called Engineering Thrive.
And there we, I proposed a metric called idea to customer, which is like, can we, instead of measuring how quickly we, you know, flow our PRs, it's.
how quickly can we flow our ideas?
So from our first planning constructs till it is in the hands of customers, how long does that take?
And I think pull request is a sort of code first artifact that's incredibly important as part of that loop, but we'll tell that broader story.
And so I think PR throughput in summary is like less a measure of overall speed and more a measure of system flow.
Still useful, just different.
So a measure of friction in the system, but we've always been very careful to make sure that people are looking at multiple metrics in tension with one another.
Good Heart's Law is in full effect.
PR throughput is an easy metric to game if that's all that we're looking at.
But the idea of the concept in manufacturing would be called concept to cash, idea to value.
I mean, that's always been the holy grail, what we really want to understand.
We're putting in more R&D units on this side.
Are we actually extracting more value for our customers and then hopefully more revenue as a result of that?
You actually hit on something that's super interesting.
You hit on lots of things that are interesting, Justin, but like one of them was...
was like particularly interesting.
Like you talk about gaming and like one of the reasons I like PR throughput is I actually think it's fairly resistant to gaming.
Like historically, if you really want to gain PR throughput, the easiest way is you just break your PRs up into a bunch of.
smaller chunks and it turns out like that's the right thing to do anyway like they're easier to review they're easier to test easier to revert they flow through the system much quicker like the act of gaming leads to what we want to to accomplish but like the newsletter you published yesterday shows that like ai is doing the exact opposite of that is doubling pr sizes and like i think that's really scary it's all that same No, I mean, that might be one of those rare times when gamification actually forces a good function in terms of breaking into more incremental PRs and more incremental delivery.
But yeah, I mean, the PR size thing, if you didn't read the newsletter, we just published it yesterday.
We did find longitudinally over a year that PR size has almost doubled.
And that should be concerning, right?
I mean, when we were engineers, you know, it was sort of part of our subtle art was the ability to pull off the same use case in as little code as possible because every line of code means a potential bug, a potential vulnerability, less portability for the code.
But then there's a psychology on top of this too, which is that like, you know, we saw this with build times.
When engineering organizations have long build times, engineers were more likely then to shove in multiple features into a single piece.
because they know that they have to wait for that build to execute.
Now, if we have an agent kind of instantly creating the code, but we still have that same delay in the build time, well, of course, we're going to try to shove more features in because we don't have to spend as much time actually even writing that code.
So multiple angles where this could be hurting us.
We've got about 15 minutes left in the webinar, and I want to make sure we've got some really good questions coming in from the live Q&A.
I want to ask a question first about token maxing.
And then I want to jump into the audience questions.
So speaking of metrics, you know, measuring token use is a new metric here trying to understand, I think, utilization.
What's your view on the role of tokens with respect to the way that we should be measuring AI impact?
Oh, my gosh.
It's just like talk about places where as an industry we have lost our minds.
Token usage, in my mind, it is like the quintessential diagnostic metric.
It tells you about cost.
It maybe tells you a little something about engagement intensity.
It tells you nothing about whether developers were able to do good work quickly and easily.
I don't even think it tells you anything about how sophisticated your AI usage is.
It's a cost metric.
And I think reporting it...
up to executives without context, it is the modern day equivalent of reporting lines of code, right?
Like it's precise, but it's meaningless.
Like as a cost metric, that is a legitimate use.
If you're spending money on tokens and we are spending increasing amounts of dollars on tokens, like, yeah, you should track tokens.
But if you're using tokens as evidence of high usage, high sophistication, you know, high better outcomes that's when you get token maxing and like that logic is backwards because tokens are about consumption and i think the right way to use tokens is again triangulating across lots of metrics and like use it as a bottom bottom layer sorry like diagnostic that helps you contextualize the movement you see in higher layers and so again like talk about id the customer which i i just mentioned if token usage goes up an idea to customer improves, it gets faster, great, you have a pretty compelling story.
If token usage goes up and outcome metrics stay flat, you have a cost problem.
And so I think tokens, they are a cost metric, often pretending to be an outcome metric, track them for what they are.
No, I when I started seeing those like token leaderboards, I mean, it made my stomach turn.
I'm like, what are we doing here?
Because you're right.
I mean, to compare it to lines of code, I think is very apt on multiple levels.
Like what we just said before about, you know, the art of good engineering, pulling off the same use case with the lowest number of lines of code that we actually have to make as part of that.
Same with tokens.
We should be able to pull off the same use case through good prompt engineering, good context engineering, good agent experience, using less tokens to achieve the same.
So yeah, no, I think we're very much aligned there as we are.
I mean, one of the things I'm so excited about doing, like researching is exploring, like what is that empirical relationship between agent experience and token costs?
And it's like, as you improve the environment around the agents, yes, ideally, like the primary driver is, can we get better outcomes through improving that human agent collaboration, but it should also improve our token efficiency.
And so we'll see that.
Let's quantify it.
100%.
And I think you brought this up before, but looking at innovation ratio, which is the way that core four looks at this.
Very stable outcome metric.
Like, are we spending more time on innovation?
And that's what we want to do, right?
If we're increasing the capacity of any engineer, is that actually translating then to that engineer being able to spend more time working on new features and creating new value?
And spoiler, we're not really seeing that.
When we put out our Q2 impact report, we've actually seen innovation ratio not fluctuating all that much.
And I absolutely talk to engineers who are like, oh, I love Cloud Code.
I can load up a spec and go play PlayStation for half an hour.
And I come back and my work is done.
And it's like, well, okay.
that has maintaining the status quo while doing less work, that's not really increasing value capabilities, right?
And so I think it's going to be really important to look at that metric.
So I think we're kind of focusing here on like, obviously we never want to hyper-focus on any single metric, but token utilization, cross-reference to innovation ratio, you know, looking at agent experience, and then using some of these proxy metrics to understand friction in the system.
If you ever find yourself reporting on a single metric in isolation, like alarm bells should go off.
Absolutely.
Okay.
I want to move to the audience Q&A.
We've got some really great questions that have come in here.
So this first one, I'm not going to read the whole thing.
I'll get to the gist of the question.
Concise and up-to-date documentation has been a goal for decades, something that few teams actually achieve.
I think that's fair.
Is there any realistic hope that teams finally document things and keep it all up to date?
And do we have any hard data indicating that that investment and effort in this is actually improving?
I mean, that's a large part of what I'm researching now.
Yes, there is hard evidence that having better documentation has led to better developer experiences.
Like an example is that from some of my previous unpublished research, I found that as developers, development teams that have higher satisfaction with documentation quality, but new higher developers on those teams.
onboard about twice as fast.
And so like it does, and like, so you would expect agents to sort of onboard twice as fast.
But I haven't proven that yet.
I sure would like to, like to your point, like it has been like one of the top pain points for developers forever.
Something like 88% of developers say that they regularly have to spend, you know, an hour or more like.
they waste an hour or more searching for what ends up being out of date documentation.
And it's just like, that's just like you're lighting time on fire.
It's so frustrating.
Everyone wants good documentation.
No one wants to write good documentation.
Now I do think what constitutes good documentation for humans will be different than what it does for AI.
Like they'll be able to infer a lot more things.
It doesn't, it's in fact, it should be a lot more succinct.
And I don't know that anyone has solved.
solved it yet, but it's something I'm looking at.
If anyone has any great research papers that they've seen or anecdotes, please put them in the chat because I'd love to dive in.
It's kind of the inverse, but one thing that we'll be publishing in the Q2, I just looked at this data today in the Q2 impact report, is we looked at individual DXi drivers and how they've shifted.
Documentation has improved the most out of any of the other DXi drivers.
So that's a qualitative indicator, but that's an interesting signal saying that maybe there is hope because we can do more automated documentation.
I think that's very good advice.
We should be splitting.
like agent memory.
So agent markdown as well as human readable documentation, we should bifurcate that, which we can do easily.
The agent should always be updating documentation when it does work.
Anyway, we should just be making sure it's part of a durable prompt or whatever that we're instructing it to update agent markdown separately, treating that more like agent memory, which to your point may look and be formatted differently than what's easy for humans to consume.
Okay, moving on.
My company has just spun up a company-wide initiative to add documentation to code bases to enable agents.
Time has been carved out from product work to make that happen.
Given we now have the time, what would you say are the two highest value areas to incorporate into the AI harnesses?
I don't know.
I'd have to think about that.
Justin, does anything hop out to you?
Only what I just said, that making sure that there's that workflow where we're updating agent markdown separately from human readable markdown.
So, and I think thinking about the substrate for agent memory is going to be a big discussion this year too.
Right now, it's just files sitting in a repo, but I think we can do better than that.
Yeah, I mean, it's just like, or just like the bloat of MD files going everywhere.
I mean, like thinking about like the validation layer, I think like the verification layer, like how do you know that it worked?
I think it's something I would focus a lot on if I invested.
huge initiative in improving context quality, documentation quality, is having a detailed plan on how do we ensure that works and what can you build into the platform itself so that you can detect regressions and things like that could potentially be interesting.
That's a really good point.
Anything we can do for better validation loops as well.
Yeah, no, that's good advice.
Let's see, we got about five minutes left here.
Several questions.
We probably won't get through them all, but we will follow up with some of our answers to the questions that we didn't get to as part of the follow-up of this webinar.
This is a spicy one.
Psychological safety.
Can we comment on the impact of AI washing of layoffs?
I don't know that I'm actually familiar with the term AI washing.
So effectively, you know, this is a statement that means like...
We're letting people off, but not really because of AI, but it's because of AI.
Okay.
No, which I agree.
I think that, you know, I've talked about this in numerous forums at Blank.
Like AI right now is really, is good at writing code.
But writing code is a fairly small part of what a developer does.
It's, you know, 14% of their day.
And like software engineering is about so much more than coding.
And agents aren't.
nearly as good at those other parts of the development sort of responsibility set.
And so I definitely agree that a lot of what we see is sort of AI washing.
And it's like, we have different monetary conditions, market conditions, all of these things that like, are complicated from an economic standpoint.
But I do think it contributes to like something you were talking about earlier, like AI brain fry.
And it's just like, I think that I worry that there is a growing burnout epidemic that we're going to have to reckon with where we feel more and more pressure to just like sprint, sprint, sprint, sprint, sprint, try things with AI.
And we are sort of getting ahead of our ability to verify and validate.
Our PRs are twice as big, but like...
Can we cognitively handle that?
Can our test systems handle that?
And it's that's all, you know, another one of these these knock on effects from we have all this pressure to go, go, go, go, go.
And I think it is going to change how we think of what our roles are is going to change our sense of identity.
And I think it's definitely going to impact our well-being.
And like maybe that's some of that pressure from from Annie's paper talking about like.
productivity and developer experience seem to be decoupling a little bit.
And we might be seeing quantifying some of that over time.
I think organizations need to be very careful.
I mean, we have plenty of compelling data now that shows us that this technology is not and may never be ready to replace human engineers for many of the reasons that you just said.
But certainly we can augment capabilities.
Certainly, you look at a company like Zapier, they're hiring more than they ever have in the history of their whole company right now because they know that they get more value out of any investment.
single engineer so i think that there's a balance right i mean if if it's a matter of um trying to populate your company with people who are really good at ai not necessarily replacing them with agents but but treating this as now a very valuable skill set I think that that's reasonable.
I think like anytime that we've had a new bit of technology and a new thing to learn, certainly people who spend time learning that are going to have an edge.
But I think that if trying to go down this path of replacing humans with agents, I don't know that we'll ever get there.
I don't think it's the right attitude at all.
I've waxed and waned even on how far does the line between roles blur.
And it's just like at one point, November came around, Claude 4-5 came out.
I'm like, oh, everyone's going to be.
developer and I like quickly backtracked on that because like some of the questions I've been seeing in chat around verification like you need someone who can recognize the failure patterns like you also need someone who has good taste on like what should we be building and what like is good enough but also How do you understand the complex systems?
How do you operate them and deploy them and manage them and maintain them?
And that requires a lot of very specialized expertise.
And I think there will always be a need for as many developers as we already have, if not significantly more.
I think we will totally induce the demand for more engineers.
We've been here before.
When COBOL was released, everybody thought we wouldn't need software developers anymore because we were going to write human language for code and business.
That didn't happen.
We ended up needing way more engineers.
It's law of induced demand.
We have a four-lane highway, too much traffic.
We build an eight-lane highway.
What do we get?
More traffic.
So I think there's an influence about a software to be written.
I do want to acknowledge, just because I believe that that is the truth.
Does it mean that non-engineering executives always interpret the data that way?
And that's why I think we see all of this volatility in sort of the job market.
I think that it is both misguided and certainly scary to be a part of.
I'm really glad you said that because, yeah, just because we know that this tech may not ever be ready to replace humans doesn't mean that a CEO, a misguided CEO, somewhere looking at the market research or whatever and the promises of 10x engineering.
could misinterpret that.
And that I agree is the real.
That's why people like us are trying to help contextualize the data.
Yes, yes, we're doing our best.
So we are at time.
We have so many great questions.
I feel like we definitely need like a follow up session here because we could get through so much more content together.
But can you tell people before we close out, you know, what's on your horizon?
What are you working on right now?
What can people expect from you in the next few months?
So I'm.
Like I have a variety of sort of papers coming out on a wide range of topics, but like the big area of focus for me is agent experience.
Like how do we define it?
What are the dimensions that sort of, you know, represent a good healthy agent experience?
You know, what is the space equivalent for agents?
And then how can we, you know, how can we measure it?
How can we improve it?
And so a lot of work on that as I look at how can we sort of optimize the human agent collaboration.
Brilliant.
I can't wait to see it.
Brian, it's always such a pleasure having these conversations.
And thanks, everybody.
I hope this was a good use of your time.
We'll be sending out follow-up.
We'll do our best to answer the other questions.
And I think we definitely need a repeat session here, too.
So thanks, everybody, for the time.
Thanks, everyone.
