# Agentic Coding Metrics and the 2026 CFO Shift

**Podcast:** The AI Native Dev - from Copilot today to AI Native Software Development tomorrow
**Published:** 2026-06-02

## Transcript

Most people, they're working interactively with one to two agents.
The most experienced engineers, they get stuck at four max.
You know, very few people get to five.
Any of us who work in cloud code every day, human attention is limited.
You can only babysit so many agents and you inevitably end up forgetting about one.
This time last year, people were debating if AI was even useful.
We were talking to customers that weren't convinced about how fast they should even move, given that AI, quote, like, just didn't work yet.
With this massive data set, the scale of it, being able to see across the entire industry what's actually happening out in the wild helps kind of ground the difference between what you're reading on X and what actually is happening at real companies.
The AI Native Dev is a podcast for developers and engineering leads at the cutting edge of AI and agentic coding.
Join your hosts Guy Pajani and me, Simon Maple, every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.
This is the AI Native Dev.
Back in November, we hosted the first ever in-person AI Native DevCon in New York.
This June 1st and 2nd, we're bringing it to London.
It's two days built for AI-nated developers and engineering teams.
One day full of hands-on workshops and one day full of practical talks on agent skills, context engineering, agent orchestration and enablement platforms and how teams are actually shipping AI in production.
Join us at the brewery in London near the Barbican for all of that, plus...
networking, parties, giveaways, and a room full of people building the future of AI native development.
You can also join us from anywhere in the world via the live stream.
As you're listening to this podcast, you get 30% off your ticket with code POD30.
Just head to ainativedevcon.io and we'll see you in London.
Hello and welcome to another episode of the AI Native Dev.
My name's Simon Maple, your host, and joining me today is Nicholas Arcolano, who is the head of AI and research at Jellyfish.
And Nick and the Jellyfish team have unveiled a huge amount of data which we're going to look through, which describes how agentic coding is done at organizations.
Some really interesting data and findings which show developers are creating and merging twice as many pull requests as they were without AI coding tools.
As well as that, we're going to be looking at the agentic barrier, how developers hit a barrier when using multiple coding agents in parallel and what it is we need to do to get beyond that.
And finally, 2026, it's the year of the CFO.
What do our engineering leaders need to bring to the conversation to make sure and show that our engineering teams are effective and productive?
Nicholas, welcome.
How are you?
I'm great.
I'm happy to be here.
This is exciting.
Awesome.
And Nicholas, tell us a little bit about, actually, tell us a little bit about Jellyfish for those who haven't heard of Jellyfish.
Yeah, so Jellyfish, we're in the AI observability space.
And in particular, we are focusing on understanding AI transformation at the organizational level.
So we're tracking what tools our customers are using and pulling in information.
Not only about the agents that are using and how they're using them, but also how those things connect to outcomes like the code they're pushing, the quality of that code, and ultimately what their business outcomes are.
Yeah, awesome.
And actually, with all that data, you have been reporting fairly regularly on a lot of this data.
And actually, I'm calling you Nick, but really, I should call you Dr.
Nick, right?
Because you have a PhD as well.
Tell us a little bit about the background behind that.
Yeah, I do.
My Slack avatar is the little, you know, the Dr.
Nick from The Simpsons.
Some folks around the office call me that.
Yeah, I mean, my background, you know, I've been doing these things.
I was in national defense research here in the U.S.
for years, you know, before data science was a thing.
And so my degree is actually in applied mathematics.
And I looked at doing kind of...
inference and signal processing on massive networks, you know, for things like cybersecurity defense.
And then realized that I could start calling myself a data scientist.
And now that title is kind of passe, right?
Did all this work to become a data scientist, and now it's AI engineer is the exciting thing.
That's it.
Always evolving, right?
Absolutely.
And I'm just looking through your CV, Nick.
And, you know, obviously you went to Harvard University, statistical signal processing, graph theory, and network analysis.
Then to MIT, Lincoln Laboratory, as a technical member of staff.
then to Runkeeper as a senior data scientist, True Motion as director of data science, and Jellyfish as head of AI and research.
So we have the right person here to talk us through about AI and data.
So first of all, tell us a little bit about the, and of course, you know, we're going to be talking about data, about trends of adoption and agentic usage through this podcast.
Tell us a little bit about the data that you collect and what you share every month.
Yeah, the data is, I mean, it's just been phenomenal, the growth of it and what we've been able to collect.
So, you know, we started for a long time, Jellyfish has been around almost a decade.
And so, you know, for a long time, we've been pulling in things like, you know, get signals, commits and pull requests, comments, things from, you know, Linear and ADO and JIRA in terms of what tasks people are doing.
It's kind of the bread and butter of the engineering machine.
And about, you know, two...
A little more than two years ago, when Copilot really started becoming Ascendant, we started pulling in data from those APIs to understand just that people were even using these tools, sort of basic usage.
And that's evolved to where now it's just a wealth of these APIs, hooks, open telemetry.
So we're able to not only understand.
you know, are people using these tools, but, you know, what's the token spend look like?
What models are they using?
And just recently, we've started getting data at kind of the agent turn level understanding, you know, things like planning and tool calls.
And so that's kind of the most exciting frontier of this data, but also just the scale of it being able to see, you know, across the entire industry.
what's actually happening out in the wild, you know, and helps you ground if you're terminally online, which I think a lot of us are these days, given the pace of things, helps kind of ground the difference between, you know, what you're reading on on X and what actually is happening at real companies.
That's it.
And that's what I love about this kind of data.
It's not it's not a sentiment data.
It's not a very, very small subset.
We're talking about tens of millions, almost 40 million PRs that were assessed in this, from what I understand.
From all of that data, what would you say are some of the highlights in terms of the one-liners that really stand out for you?
I mean, one of the biggest ones was it's amazing how fast it's changed.
And you'll remember, Simon, you know, this time last year, people were debating if AI was even useful or at least outside of maybe the bubble of us that had already bought into the utility of this.
But we were talking to customers that...
weren't convinced about how fast they should even move, given that AI, quote, like, just didn't work yet.
And we know what happened at the end of last year.
I mean, there's been a couple of sea changes, obviously, but, you know, the models that came out in the fall and kind of how that advanced agentic engineering.
So, you know, we find ourselves now with this massive data set, we've really cemented the understanding that.
There are real raw coding games, and we see them out in the wild.
And we find ourselves now, with this massive data set, we've really cemented the understanding that there are real raw coding games, and we see them out in the wild.
And these fancy IDEs, they're using Cursor and Copilot.
Those things are increasingly agentic.
And without changing other material things about your workflow, 2X is about kind of what you can do in terms of just raw code throughput.
And so the 2x there is the increase in the number of pull requests that are created, right?
That's absolutely right.
And are there types of applications or types of, I guess, projects whereby you see those increase faster or types of projects or applications whereby you actually see AI maybe not as well used?
Yeah, that's a fascinating question, and we do see that out in the wild.
Some of it matches what you would expect, which is smaller code bases, newer code bases.
We see language differences.
Those things, you know, are faster with AI.
So things, you know, the languages that we all kind of experience are more AI friendly.
Things like Python, TypeScript, you know, things that are very heavy in, you know, configuration, Markdown, YAML.
You know, those things move faster.
Things that move slower are older code bases, big, messy distributed code bases.
We have results that show you essentially.
get little to no gains due to AI, you know, as your code base becomes very distributed, where there's just a lot of human work involved, both in terms of mapping together the context of how all this code relates.
You know, I think we've, you know, we've all been in the situation and know the people who...
You know, you've got the senior engineers who know where all the bodies are buried.
And if you actually want to get deployed, they're like, well, how, you know, what do I have to do?
And how does this relate to that?
And it's just knowing like, oh, you have to change the code in these 10 places across these repositories.
That stuff isn't operationalized yet for those big, messy, sprawling code bases.
And you just don't see the gains due to AI there because the agents don't know what to do.
It's still heavily human in the loop.
Yeah, yeah.
Super interesting.
And I think it'll probably.
I don't know whether there's enough information there, enough data there to say, well, actually, this is the style of application whereby AI will really, really help you, or whether that's just, you know, natural because people want to use it in certain situations because they're more familiar themselves.
They understand knock-on effects and those types of things.
But it'd be fascinating to kind of like dig into that.
I guess from the pull request point of view, When we talk about the 2x, is that 2x merged pull requests or raised pull requests?
Those are merged pull requests.
So it's 2x merged pull requests.
The fascinating thing and the thing that's kind of mind boggling is we're starting to see, but for a long time throughout the past year, you know, six months even.
you weren't seeing downstream increases.
So you weren't seeing increases to the number of features shipped.
We have a view into things like not just issues, but epics, and those things roll up to initiatives.
So our customers at Jellyfish do a lot of work to use our platform to map those things into the business units that they care about, product lines, customers.
And we weren't seeing movement on those things, and we're starting to.
there's more capacity, but where's all that capacity going is kind of this mystery that we've been trying to untangle for the past several quarters.
Yeah, interesting.
I'd love to dig a little bit deeper into the number of pull requests here, because one thing that's very curious, and I feel this from my previous role at Sneak, where Sneak was very, very good at identifying issues, raising those pull requests, and...
Those pull requests will sometimes get merged, but a lot of the time they won't and they're there and they're valid pull requests, but there's extra effort that is required to merge them.
There was some fascinating data here about merge rates of AI generated pull requests versus human pull requests.
Tell us a little bit about that.
Yeah, so I was just looking at that recently and I published some results there where You know, the average for Q1 for humans was about 80% of pull requests that were open ultimately got merged.
And the other 20%, they either stay open or they get closed without being merged.
There's a variety of reasons that that happens, right?
It might not have met quality standards.
You might have decided it's not a good idea.
What we see with agentic PRs.
that is 60-40 instead of 80-20.
So you're talking about double the amount of PRs that are kind of dying on the vine, so to speak there.
And, you know, we're digging into what are the reasons for that.
But anecdotally, when you talk to people, some of that is, you know, workflow differences.
So, you know, we have customers we talk to, and I'm sure you know people who they're doing these things proactively.
A customer request comes in.
You just tell an agent to code a fix, you know, so it's just there.
And then you decide later if you even want that fix.
We do have, you have the workflow where people might do things, you know, two, three, five different ways.
You're not sure the architecture you want.
You're not sure the way you want the feature to work.
So you, you know, do a bunch of different candidates and then you pick one.
So that's an intentional kind of throwaway work there.
And then there's also the version we hear a lot about where, you know, people see these things on the backlog, they see changes, and they sort of, you know, to use the term that's falling out of fashion, vibe code, a fix.
And then people realize it's more complicated than they thought.
It doesn't meet the quality needs.
The senior engineer steps in and says, no, no, no, no.
We didn't fix this for a reason.
There's your unearthing demons, right?
There's deep technical issues buried under this seemingly simple thing.
So all of these kind of compound to give you that doubling of the rate.
And the reason I think it's so interesting is because this gets at the question we all want to know, which is, Ultimately, what does good agentic engineering look like and what is it going to cost?
And so is it, you know, how much waste, you know, overhead is there going to be when you're doing highly autonomous things, things with teams of agents?
You know, it's not a simple scaling of what people are doing today with, you know, one or two interactive agents.
Really interesting.
I actually assumed incorrectly that this was as a result of maybe more autonomy on the agent's point of view to, I'm going to, you know, do some stuff and create a pull request based on a trigger I found versus a human actually, you know, initiating these requests and then choosing to pick one.
So it sounds like, not to paraphrase, but it sounds like, is this a flaw?
in our workflow whereby a GitHub or something like that, you know, a Git repository just isn't cut out for agentic development whereby we need a layer there that should expect, you know, multiple...
versions of a fix and allowing us to pick one.
And actually what we're doing is we're creating a workaround where rather than send one pull request, which is the fix, we're actually sending multiple and throwing some away.
It feels like we're trying to use the technology here almost as a means to what we can do with agentic development right now.
Yeah, I completely agree with that.
We know that, you know, the whole world of software is littered with tools that maybe aren't being used for the purpose that they were designed for.
And, you know, some of those persist and we just use them in these weird new ways forever.
And some of them don't.
So I think history will tell whether, you know, my my kids are doing agentic engineering and asking, why do we do all this weird get stuff?
And we're like, well, that's just it's been around for decades and it's not going away.
But it's not it's not the right.
The question of, you know, what is what is the pull request look like is kind of the the core unit of software engineering, I think is a fascinating one, because, you know, right now it's so tied to value in the shipping.
And we see that whole SDLC collapsing in new ways that, you know, I think it doesn't necessarily make sense anymore in a truly agentic world.
Yeah.
Yeah, super interesting.
It's amazing how many DIY tasks or home improvement tasks I can do with a hammer.
My wife will my wife will swear at me many times.
So let's move a little bit on to a new topic here talking about AI adoption within organizations.
So you've said in the report, AI adoption has reached a median of 71 percent of developer time.
That seems that seems a lot.
Tell us a little bit about how that data is generated, how you how you're learning from from, you know, AI adoption statistics here.
You know, is this something that is something across our entire industry or within certain groups?
Yeah, in terms of adoption, it's interesting.
I think for some of us, it seems like it's like, why are we even talking about this?
All of us use this all the time.
And, you know, at Jellyfish, we have the benefit of just seeing such a broad swath of companies.
And so, you know, all of our customers, they understand that they need to be doing.
AI coding, they need to be doing agentic coding.
And the basic adoption barriers are still very real for a lot of companies.
Most of them have gotten over the hurdle of just the basic, you know, security enablement things of can I even get licenses to these things?
But some of them are still just measuring, you know, do people even use these?
And if they do, let's say there's a mandate to use them, are they just checking in in a performative way or are they actually integrating them into the workflow?
You know, I think it's a vanishing minority of people who, you know, are highly resistant to these tools and maybe they're logging on and using them but really don't want to.
What's much more common and so that 71%, that's weekly active users across, you know, the entire.
developer base that we're tracking about, you know, 250,000 developers in total at this point.
And so, you know, the median is 71%.
P90 is 90%.
So 90th percentile is people, you know, using it at...
90% of them, essentially all the time, except when they're on vacation or maybe stuck in meetings all day.
Even then, we're using them in meetings all day.
What's the typical background of these 250,000?
Is there an average developer here?
Are they working?
Are they more open source?
Are they more freelancing, those types of things?
Yeah, our customers, so they're paying us to track their AI transformation and how that maps to the business outcomes.
So they tend to be...
They're companies that either build and sell software or software is part of how they do business, but they might be, say, an apparel company, but they still have 200 engineers because they have operations and they have a digital presence.
So you need jellyfish when you get to be...
a couple dozen, you know, tens of engineers.
And then we have customers up through tens of thousands of engineers.
You know, what we don't see are the, you know, five people kind of, you know, in the valley that are, that have started an AI native company yesterday.
They're, they're not using us, you know, cause they just got started.
They haven't run into the problems of, of needing the level of observability.
And we hope that they do eventually.
But those are the types of companies.
We don't see a lot of open source models.
So these are a lot of companies, you know, they're buying kind of big enterprise tools and licenses.
So they're using Copilot, Cursor, and Cloud Code.
Cursor and Cloud Code are dominant.
Cloud Code, as you might expect, has just been ascendant over the past six months.
And they're using those models both through those providers and they're using...
Bedrock models.
And so it'll be exciting now to see now that you can use the open AI models on Amazon Bedrock.
I think we'll see a lot more uptake of those.
We have a lot of customers that use Bedrock.
Yeah.
And I'd love to go into a per tool, actually, analysis to kind of understand if there are any trends per kind of like agentic coding platform.
But I'd love to understand in terms of the weekly active use, and this is something that I'm super interested in, the depth.
of use of these tools, how much would you say you have data that show how much people are relying on agentive development in their roles as well as being a weekly active user?
Yeah, that's a great question.
And there's a couple of different lenses on this.
So one basic lens.
So you've got, you know, is this person even touching these tools on a regular basis?
That's, you know, the very basic table stakes.
And then we look at, you know, do people have a reliable habit?
Are they using them repeatedly, you know, day after day, week after week?
And so one benchmark we said is, are you using AI tools regularly three or more times a week?
And that persistent, frequent, active usage, that's what correlates with and how we are able to see companies that are able to drive up persistent usage across their users.
they see those productivity gains that I was talking about, that 2x.
And then the level beyond that is autonomy and kind of agentic workflows.
And that's where we see the, we look at things like what percentage of your PRs are being generated in an autonomous fashion.
So an agent is able to take a spec and actually do all the work, open the pull request with minimal human interaction.
And that's an area where we've seen kind of fascinating growth and really different stories.
So the, you know, the elite companies, the 90th percentile that we look at, those folks, they last month, they were they crossed 20%.
This month, we're crunching the numbers, it looks like it's closer to 30%.
But it's basically been exponential growth.
You know, the and they were at, you know, say 2% about a year ago, the median company That company is just past 2% recently.
They're at about 2.5%, 3% now.
So it's really a story of the leaders here, the people that are figuring out this agentic development.
They're kind of running away with it in a lot of ways.
And everyone else is still kind of figuring out how to get these workflows repeatable, scaled.
They might have a small pilot team that can do these things, but they don't have it.
you know, kind of scale to the whole organization.
Right, right.
And I think that's very classic of rolling out any kind of technology.
So it's similar to how we've seen before.
I'd love to talk about the top X percent, maybe those who are the very heavy users of agentic tooling.
And it kind of like leans us into something that you refer to as the agentic barrier concept.
So these more elite, very experienced agentic developers who are using many different agents in parallel, tell us what the agentic barrier is and when they reach it.
Yes.
So this is something that we observe.
So this is work we did in my team and a colleague of mine, Tomas Pardenas, you know, he's gone deep in understanding these agentic workflows.
And what we see is, you know, most people there, they're kind of working interactively with one to two agents.
And we see the most experienced engineers, if they have to do kind of any interactivity with these agents, they get stuck at you know, four agents max, you know, very few people get to five.
And it kind of jives with what, you know, any of us who've worked with these, you know, who work in cloud code every day, like I do, you just, you can only, you know, human attention is limited.
You can only babysit so many agents and you inevitably end up.
forgetting about one and you're like, oh, I forgot that was even a thing.
You know, you missed the push notification to go click on that terminal window and you forget to go back.
And so we find that even at that level of four concurrent agents, you end up spending 80% of your time just focused on one.
You know, interactivity just has its limits.
And so, you know, to break beyond that barrier, you have to get to a fully autonomous mode.
You have to be able to hand off work in its entirety.
You know, if you have to nudge the agent along, things just kind of grind to a halt.
And the, you know, the total concurrency is just kind of has a hard limit to it based on human attention.
Yeah.
And I guess that's almost like allowing instead of the human being the orchestrator, you're just enabling.
and another agent to do their orchestration, which is essentially handing the whole thing off.
Do you feel like there's a better way that we can work with multiple agents, or do you feel like this is a technology issue?
There's, what makes this a hard problem to answer is there's clearly technology issues, and we've seen advances in agents and interfaces.
You just talked about pull requests.
You know, lots of us agree that, you know, the tools that we're using to to manage code, to manage agents, to inspect what's going on are limited.
And I think all of the players here in this space are, you know, all of them would love to invent what the new interface, the new IDE for agent development is, right?
That would be, you know, a trillion dollar, you know, invention, certainly.
But I'm also, you know, I have my own skepticisms and part of what I'm trying to understand, you know, dissecting all this data about the SDLC is, you know, what are the limits in terms of, you know, actual business and product outcomes?
So what I mean by that is, you know, there's governors, there's a speed of light associated with how fast you can make business decisions, how fast you can get feedback from the market, how fast you can enable your go-to-market, how fast you can change your product.
I hear product leaders talking about starting to deal with user fatigue of products changing too fast.
And you don't even have time to get feedback on whether these features are better because people don't have time to use them and they're becoming increasingly frustrated with how fast the products are changing.
So understanding when are these throttles on our ability to actually build products ultimately then throttles on how fast it even makes sense to build software and how interactive you really do want it to be.
Yeah, it's a really good point because I think, we've got to insistently remind ourselves that our users are still our users.
And, you know, we're not, well, we're sometimes, we're sometimes, you know, using technology or agents as our users, as our customers.
But, you know, the majority of the time, it's going to be a user that's going to be using our software in an end user point of view.
So, yeah, I guess that's one style where actually slowing down a little bit and having that more interactive flow.
can be better because we can introduce more of our thoughts and our takes.
Are there other ways in which autonomy or autonomous agents is actually the wrong call?
And how does that adjust things like team sizes, how we interact with other folks on our teams?
Well, certainly, you know, we see a lot of customers who are, you know, experimenting with and thinking about these smaller team sizes and the ratio of product folks to builders.
You can just do so much more with an engineer now than you used to be able to.
And so, you know, I am certainly of the belief that a good product leader, someone with good, you know, user and business sensibility is just worth their weight in gold or tokens or whatever the most valuable commodity of the moment is.
So, you know, I think I think that's a thing that is is changing.
And there's a real question I find myself asking the question when I sit down to do a thing, you know, how much should I brainstorm with the agent and figure out exactly what I want and then hand that thing off in its entirety?
Or how much should I kind of work through it and have the agent help me understand what I want?
I think I'm in a common position that a lot of developers and creators are that I sit down and even if I think I know what I want, I'm quickly disabused of that notion that I have thought this through in any meaningful way.
And engineering, so much of it just depends on these really...
kind of sometimes high stakes, you know, high tech decisions about things like architecture, right?
Trade-offs in the gray area, things, those are really the meat of true engineering and true product development.
And I think we're all trying to figure out how much of that comes through friction and exploration and banging your head against things.
And if you outsource too much of that to the agent, do you lose the opportunity to have those insights because you're not in there?
Hey everyone, hope you're enjoying the episode so far.
Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development.
Whether that's talking about the latest tools, the most efficient workflows, or defining best practices.
But for whatever reason, many of you have yet to subscribe to the channel.
If you're enjoying the podcast and want us to continue to bring you the very best content, please do us a favor and hit that subscribe button.
It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you.
All right, back to the episode.
I'd love to talk a little bit about, I guess, one of the differences from 2025 to 26.
You mentioned 2026 is the year the CFO gets involved, which is a very nice way of putting it.
I guess in 2025, people were much more encouraged to try everything you can.
Let's learn.
much, much more about what is out there and what suits us well.
And the only way to really do that because of, you know, everyone having different cultures within their organization is to try it themselves and see what works for them.
I guess now people are a little bit more cautious in terms of, you know, spend and trying to think, how do we make the most out of our budget with AI?
We still want to spend in AI, but we're much more thoughtful about our budget.
I guess from the point of view of data, do you see a difference in 2025 to 2026 in terms of how much variety there is in tools that are being used?
Are people focusing more today on a subset of tools or are they still, do you see people still growing in their exploration?
We see, so we definitely have seen and we advise people to, at this point, so many of the tools have gotten so good that, you know, the risk is kind of, you know, getting analysis paralysis and trying to find the exact right agent, the right workflow.
Should I do spectrum development or not?
Should I use codex or cloud code?
You know.
For the vast majority of cases, picking a horse and riding it and just getting good at exercising that, building that muscle is the right move.
So we do see a lot of that.
The big difference is the kind of move to scaling this year.
So a lot of folks, there was budget and urgency to just explore and try things out.
And now as companies try to scale, you know, we've seen how much more complicated, you know, how much more complex agents have gotten, how much more capable they are, but also how many more tokens they burn.
You know, the models are beefier.
You know, the agents are doing more turns.
They're doing more code exploration.
So token costs have, you know, token consumption has just skyrocketed.
And, you know, I think that the tension that engineering leaders are feeling is there's still very much a, you know, it's like the token maxing conversation.
There's still very much a perception that, you know, adoption and regular use of these things needs to be driven, needs to be supported.
And so if you're not using this stuff enough, then you're not going to get there.
So you've got to be burning enough fuel to show that you can escape.
You know, you can achieve some escape velocity, but you have to show your receipts.
You know, we're at the point where engineering leaders also have to explain where are all the tokens going?
Because it is a lot of money.
And what does that look like?
I guess in terms of if the CFO says that we've spent all this money, we need to see some level of either efficiency or productivity that our teams can build faster, deliver faster.
What does that look like?
from an engineering manager's point of view, what should they produce back that the CFO would be happy with?
Yeah, it's a real mess, Simon.
I mean, the challenges.
That isn't the answer, I presume, that you give back to the CFO, right?
No, no, no.
No, what's fascinating about this right now is that, you know, so there are good metrics.
So you can look at raw token spend, right?
That's sort of a basic signal of, you know, are we, like you're, Like your heart rate.
Having a heart rate is a good signal you're alive.
And then you start to ask more sophisticated questions, which is, you know, what is the token spend per PR, per feature shipped, token spend per outcome?
And that gives you a sense of what you're actually spending.
The reason I say it's a mess is because, you know, What businesses really care about are things like revenue, right?
They care about things that actually change, you know, the trajectory of the business.
And companies generally haven't adapted to think about what does it mean to essentially have infinite capacity to be able to exchange.
I mean, on kind of the coarsest level, we've entered a world and are increasingly in a world where you can just spend infinite money to get things built.
You can hire infinite robot contractors, so to speak.
And companies aren't designed to reason in an infinite capacity world.
They're designed to think about, this is the headcount we have.
This is how many things we can build.
What's going to maximize our business?
So that feedback loop of what would actually matter, that's still being built.
And I think that's hard for people to reason about where should you apply maximum leverage.
It's funny.
I always tell my team, To focus on outcomes, not output.
So when we look at the data here and we see a 2x in pull request merge rate, for me that's output, right?
It's like, you know, we're getting PRs created, PRs merged.
In terms of the outcome, is it good code?
Is it quality code?
Is that code being patched?
You know, is the second PR the fixing of the first PR?
How do we actually identify if this is truly...
you know, being faster, making us, putting us in a better place than we would have been without AI coding.
Yeah, those are important outcomes for the engineering machine.
And we see lots of folks, you know, concerned about quality.
And right now there aren't huge smoking guns in terms of quality.
You know, I think the main reason for that, that we haven't seen, you know, quality go off the rails is that engineering teams by and larger.
responsible stewards of quality.
So they're not merging bad code.
You see that difference in merge rates.
So you do see some, you know, upticks in reverts.
It's, you know, it can be hard to analyze because we also have seen a growth in fixing forward.
So as teams get faster, they may not even file bugs or revert code.
They may just apply a fixed forward.
So it can look like just more building and not necessarily fixing a problem.
But all of those things, you know, those, again, are engineering outputs.
They're still not necessarily business outputs.
So, you know, did you build product or did you just kind of tinker with things that maybe nobody else in the org cares about but engineering?
And even if you built a new product, you know, I think, you know, we have go-to-market organizations and things that, you know, if I could.
wave a magic wand and build five times the products tomorrow for Jellyfish, could we sell five times the products?
Probably not.
We would need, you know, a whole new type of AI enablement and our go-to-market org to accelerate that team to a place where they can even capture all of that value.
So that's what I mean by, you know, we're in a world where we've invented these jet engines and we're still kind of putting them into the...
cars we used to have, we haven't designed the rest of the vehicle to accommodate this one thing that we've massively accelerated, which is code generation.
Yeah.
And as we talk about becoming more effective, becoming more productive, and scaling that across an organization, I guess one important role or area that perhaps we've not looked at as much is more around the the ai enabler uh or you know maybe it's part of the platform team the developer experience team the the group that essentially encourages and and helps teams and organizations roll out uh ai adoption across their organization do you have experience or or any any wisdom you can kind of like share in your in your discussions about how organizations are getting good at these types of rollouts and gaining adoption through sharing within that organization?
It's hugely important, Simon, because we see this kind of uncanny valley where small teams, because they're small and also because they tend to be younger and have younger code bases, the company itself just doesn't have all the baggage of an older company.
They can move fast because they're small.
And we actually see very large companies that have the types of teams you're describing.
They are able to make big investments.
And the ones who are the most successful are, you know, they're moving with a purpose and with deliberation to enable folks.
They understand that, you know.
You need to invest in the tools.
You need to invest in the context engineering.
You need to invest in training.
This just doesn't happen for free.
And where people struggle is in the middle where they don't have those teams.
They may not even have a head of developer experience.
That person certainly doesn't have a whole team of people that can develop tooling.
And when you kind of leave every team to figure this out on their own, it can be really challenging.
So, you know, what success looks like in my experience, what I've seen is, you know, putting dedicated resources to it, making it someone or multiple people's jobs, full-time jobs to figure out how to make this transformation in the org, you know, making clear investments with money, with training, with time, going slow to go fast, you know.
And then the thing we talked about before, which is just picking some things and not getting analysis paralysis or stressing about the fact that certainly tomorrow.
a new thing's going to come out that makes the thing you just did not the optimal thing to do.
But just building those muscles, you know, it comes down to continuous learning and continuous evolution at the end of the day.
If you can't build that muscle in your org to just evolve continuously, we're all on this treadmill for a while, for the duration, right?
It's just going to keep accelerating and you're just going to have to keep learning new tools and evolving your org.
So that's what it looks like.
Very interesting.
Keep going on that thought process of giving advice to engineering leads and organizations.
What would you say is the biggest thing that folks are getting wrong or the biggest misconception that engineering leads have today around AI?
I think one big misconception is what we were just talking about would be the flip side of assuming there's some silver bullet.
If I just give these tools to people, they'll just magically work.
I think part of that misconception is we've kind of done the easy part already in the sense that, you know, we gave a lot of people fancier IDEs and there were real gains associated with that.
You know, fancier autocomplete, they sort of, you know, they bootstrapped a lot of coding and there was some training involved.
But by and large, there have been real gains associated with getting a co-pilot or a cursor in the hands of developers, but then kind of leaving everything else in the process the same.
So, you know, the misconception is the understanding that to get to the next level, to get to the promise of, you know, what you and I believe AI native development is really going to look like.
Those are big cultural changes, the big skill changes, the big architecture changes.
They are much bigger rocks to move and require much bigger investments.
The payoff is going to be massive.
You know, you and I certainly believe that.
But I think, you know.
Some people aren't prepared for how different that is than what we've done so far at a lot of companies.
Yeah, it's almost like questioning everything, right?
It's like, don't assume that just because you used it the last 10 years, just fine.
It's the right way of doing it with AI.
And that said, actually, there's probably a lot of things that we should very intentionally keep that AI sometimes makes it easy to drop.
A lot of the typical best practices and good hygiene processes.
of good software development uh you know things like your you know your usual code reviews and the best practices your style guides potentially not so much but the good practices in your workflow a lot of that we need to maintain because it's sometimes too easy to to vibe code and and throw something and it's making that assumption oh actually yeah i'm sure there's no security issues here and and so forth Yeah, I mean, these are, you know, people are having these kind of soul searching arguments.
We've had an argument internally at Jellyfish that a lot of folks have had, which is, you know, we've done the two person kind of code review thing of, you know, if I open a pull request, a different person needs to approve it.
I can't approve my own pull request for production.
If an agent wrote the code, is that a different person?
Can I review code that a robot wrote that I never read?
It's kind of, you know, logically it kind of makes sense in trying to decide, you know, is the, you know, the throttling of, I mean, reviews are becoming a bottleneck, you know, is that worth the value of having two different eyes on the code?
And it, you know, I think about, I love the way you phrased all that about, you know, what things do we want to hold on to?
One of my pet peeves or things that I'm trying to understand about the current conversation is this obsession with agent coherence and how long an agent can run on its own is such a weird metric.
And I understand why we're grading and benchmarking agents on this.
But then when you translate that to the real world, I always use the analogy.
You know, if you think of your team, is the best engineer the one who goes off the longest without talking to you into the code cave and asking any questions at all?
Like we, you know, we tend to complain about those folks who are, who don't know when to come up for air.
So it gets back to our kind of interactivity question of, you know, feedback and communication is a core part of a really well-functioning engineering team.
And what does that look like in an agentic world?
When is that agents communicating to each other?
And when is it they need to be communicating to us and having them go off on their own for a very long time is just is a is an anti pattern.
It's a dysfunction.
Yeah.
Yeah.
Super interesting.
Nick, where can people go to to keep on track, keep keep on on top of all the all the great the great work you're doing with your reports?
keeping up with the data, that kind of thing.
So if you go to jellyfish.co slash AI engineering trends, we are publishing updates to our benchmarks monthly.
So, you know, we have a history of previous months, things around adoption and impact, quality, you know, agentic growth and agentic engineering.
So we have all those things and we're adding new stuff every month.
So we just added some token insights.
We're going to add some more in the next cut.
So we just can't keep up with all the exciting things that we can see in this data.
you know, not enough hours in the day or tokens in the world to understand everything.
I would love to understand about what's actually happening in the real world right now.
Amazing.
Well, great shout.
Let's keep up to date with that and looking forward to the upcoming reports.
Thank you very much, Nick.
It's been a pleasure.
Yeah.
Thank you so much, Simon.
This was fun.
Absolutely.
And thanks everyone for listening and be sure to tune into the next episode.
Bye for now.
The AI Native Dev is brought to you by TESL.
the package manager for skills and context.
Your hosts are Guy Pajani and me, Simon Maple.
Our producer is Tom Dowler.
The AI Native Dev is not just a podcast, it's a community, and we host monthly meetups at the Tesla offices in central London.
Visit tesl.io forward slash community to learn more, and I hope to see you there.
