# Scaling Agentic Development: Governance, Metrics, and Workflow Shifts

**Podcast:** The AI Native Dev - from Copilot today to AI Native Software Development tomorrow
**Published:** 2026-06-25

## Transcript

Now that it's agents doing the work and tokens and the cost of those, everyone's going to be really keen on making sure the software development lifecycle was as effective and as efficient as it possibly could be.
But when it was humans, it was like, oh, well, Timmy's just lazy and that's why stuff isn't getting delivered.
And they didn't focus enough on process and value stream mapping and all those important things.
Now when it's agents, it's like, well, that's going to cost us money and we can easily see that it's costing us money.
Most people know me from the DevOps transformation and the Cloud Native.
transformation and I love every new chaotic period in the industry because that's where the learning is happening and right now that is EI and coding.
The way I think about it is I go back to pre-gen AI and I imagine I had all the manpower and all the time in the world and basically with that mindset we say okay now you have gen AI, you have agents, there is no excuse not to do everything you ever thought you want somebody to do and then automate it so it happens by CI automatically.
The AI Native Dev is a podcast for developers and engineering leads at the cutting edge of AI and agentic coding.
Join your hosts, Guy Pagiani, and me, Simon Maple, every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.
This is the AI Native Dev.
Hey everyone, hope you're enjoying the episode so far.
Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development.
Whether that's talking about the latest tools, the most efficient workflows or defining best practices.
But for whatever reason, many of you have yet to subscribe to the channel.
If you're enjoying the podcast and want us to continue to bring you the very best content, please do us a favor and hit that subscribe button.
It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you.
All right, back to the episode.
Hey there, Simon Maple here.
Welcome to another episode of the AI Native Dev.
This week, we're sharing another fascinating conversation that we had with some of our AI Native DevCon speakers, this time on the topic of AI enablement and what enterprises need to get up to speed with agentic development.
I sat down with the godfather of DevOps, Patrick Dubois, Autonomy AI's, Tamuz Dubnov, and Daniel Jones from Resync.
What I think you'll agree was a really intriguing discussion.
And what I love is the way that some discussions, you very often get question, answer, question, answer.
But these folks were very, very passionate.
They really added to each other's points.
And it was a true conversation.
I hope you enjoy.
Hello, Simon Maple here, and I am at the AI native DevCon.
2026 in london and we've got a wonderful panel here to talk about ai enablement within organizations how we can be most successful at scaling agentic development across our organizations.
I've got a wonderful panel, as I mentioned.
We have Daniel Jones, DJ, head of product at Resync.
We have Tammuz Dubnov, who is the co-founder and CTO of Autonomy AI.
And we have Patrick Dubois, who is the dev role at TESOL, the DevOps overlord and godfather of all things pipeline and workflows and DevOps.
Welcome, everyone, to the panel.
How are you all doing?
Not too bad, given that there was a party on a boat yesterday.
It was.
Yes, and I'm just about functioning, so that's good.
Well, that's always a benefit.
Even with good weather in England, right?
Yeah, yeah, even with good weather.
On a boat, on the Thames, and it was sunny.
Who knew?
Who knew that was a possibility?
So why don't we start off with you, DJ, and let's go left to right, and tell us a little bit more about yourself, your role, and maybe a little bit about Resync.
Sure.
So I am head of products at Resync, which means productizing our services.
We are an AI native transformation consultancy.
We were big in, the co-founders were big into cloud native transformation, wrote the O'Reilly cloud native transformation book.
And what we see is that organizations are repeating the same kind of mistakes when trying to adopt a disruptive technology of...
just plopping the technology in, not changing anything about how they work, and being surprised that they're not going faster.
So we offer various services around that, and I do a lot of the training, agentic coding training.
Awesome.
Thames?
So Thames is co-founder and CTO at Autonomy AI.
We built an operating system on top of your code base to enable an easy platform for non-technical or semi-technical individuals to actually meaningfully contribute to the code base in the organizations, something that can do their visual iterations and actually get stuff.
ready, do backlog items, do new items, put everything inside a Brownfield codebase that ends up in a meaningful PR that can actually be merged.
Patrick?
Patrick Douai.
I am from Belgium and I have gray hair, so I've lived through a bunch of transformations in my life.
Most people know me from the DevOps transformation and the Cloud Native transformation.
And I love every new chaotic period in the industry because that's where the learning is happening.
And right now that is AI and coding.
And so I'm looking not just at the tools, but also where kind of the new organizational patterns are happening as this field is forming in the industry.
So that's kind of my jam right now.
And every time there's a chaotic transformation, the hair gets grayer, right?
Yeah, it's hard to know whether it's the cause or the effect.
Yeah, I know.
Brilliant, brilliant.
Around this table, we're going to be chatting about all things AI enablement.
Let's kick off with an interesting question about when we want to scale AI, agentic development across our organization in a way where we can control it, it is something we can govern, it is something we can do in a meaningful and controlled way.
Who owns that?
Which team, which individual, who's responsible for making sure that happens well?
It's a good question and one that I think a lot of people are grappling with and I should imagine that we'll probably end up with new departments, new teams, new named units and movements forming.
I'm working with one customer who happens to have a developer experience team which is kind of overlaps with the platform team and so there's an amount of kind of, it makes sense for those folks to be doing things like building and distributing and assuring the quality of skills.
for example.
Platform teams are kind of interesting in that they're there to help developers by presenting a raised level of abstraction, reduce cognitive load.
And you can imagine that making sure that agentic coding goes well and context is managed and available to agents would fit into that group of people's responsibility.
But I've seen other customers as well who have a DevOps team, which probably shouldn't get too deep into that.
whether it's an anti-pattern or not.
But I've seen those folks, like the people that get lumbered with like, you're doing the CICD pipelines for people.
Like those people also being kind of made responsible for these things.
Patrick, I'd love to go to you there.
How many hours of sleep did you miss when the DevOps role came about?
None.
None?
None.
No, it's quite easy.
The big thing is that every new change in an organization usually has like a small incubation kickstarting team.
We had the Agile team.
We don't talk about an Agile team anymore.
Everybody's doing Agile.
We had the DevOps team.
That kind of is just a way of scaling things out.
You need that one team that gets duplicated in multiple teams like, hey, can I repeat this?
Can I repeat this?
Then I think the pattern is indeed right now is it falls back to a central team kind of nurturing all the other teams in a way.
I think the challenge, what you were saying, like, yeah, it might end up in the developer experience platform team as a role.
Their challenge right now is they're still often focused on, oh, we'll provide the infrastructure.
They're not the AI-savvy tech coders.
So that's where there's now a little bit of a vacuum in kind of the companies where like, okay, you know, they're not the best fit, but they do control the spend, they have the models, they have the gateways.
So you can see their stack kind of expanding.
they will stay in one group as a platform that's uncertain.
What you might see is that you're going to have the cloud platform team, the AI platform team, the data platform team, and they're all kind of platform supporting this.
So it's not because there's like one platform group that they have to do it all.
And that's the same thing with like feature teams.
They can't do it all because they have a specific focus.
Small companies, yes, same team, that alliance.
Bigger companies, they might have different skills that they're going for.
Anyway, that's kind of how I look at it.
Tell me, as a CTO, you've got your engineering team there.
They want to use AI as much as possible.
They don't want to be slowed down through extra governance and things like that.
Where's the balance between having a team wanting to...
add governance, add ways of working, showing a golden path, versus at this stage of AI's maturity, immaturity, allowing people to kind of find their own paths and understand what's good through developers playing around and trying different things.
Sure.
So for us, SpeedUp has been getting the whole team aligned.
Each person finds a new skill that's important in their workflow in some manner.
And then we have a way to share with all of them and have everybody else they use it.
It self-learns.
But the mechanism that's really kind of sped up our R&D is getting to focus on the stuff that really matters.
And we get that because we offload the stuff that doesn't need an engineer to our product team and our designer team.
So part of the game for us has actually been, well, we have builders that are engineers, builders that are not engineer.
How do we manage the non-technical opening to PRs?
And how do we manage the technical side getting those PRs?
And that kind of reshuffle actually, that's what has really helped us get through that AI native speedup.
Because you have to do change management internally, both on the receiving end and the new builders.
And suddenly when that does click, that's where you see everything running forward and you see the whole team kind of playing together in harmony.
I think that's interesting because, you know, when you called...
in the past, Dev and Ops bridging two worlds, like developers and ops coming together.
I think the term that comes up a lot now is the AI product engineer.
They're in both worlds trying to cross the silo and the bridge.
You were working on how to align that better as a field.
I think that's fascinating in a way that when the toil or the daily things are being more and more done by AI, we can start redirecting them.
It's not about like, How are we building the thing right?
Are we building the right thing?
That's where the product focus comes stronger again.
I think I like how these roles are blending.
It's not for everybody, I guess, but it is a good alignment to have in an organization.
We see that there's the span of work that needs to be done, and frankly, there's the span of people that care for different parts of the work.
There's the stuff that the designers and the PMs really care for, and honestly, that's not the stuff that the developers care for.
So the fact we can say, great, well, those builders are now doing it.
Hey, developers, you gave our architecture, scaling, the sandbox environments and how they live.
You can focus on that because all this, let's say, annoying work you used to have to do.
Well, now the builders that care about them, they also have the authority to make those decisions.
They're the ones doing it.
And hey, you can focus on the stuff that you care about and making them more agent, tech-friendly and faster developer.
And that's just happening in our team organically.
They're already there.
They're already trying to speed up.
And they're already figuring it out on their own.
And for me, the focus as CTO is enabling the site that isn't pushing or wasn't pushing as much, which is the non-technicals.
And when you get them up to speed, suddenly you get the full speed up and you get people working on the stuff to take care about.
And then that's the stuff that they're fastest on.
Yeah.
And speed is the most interesting thing here because I think, you know, when we think about a year ago, two years ago, Gosh, I'm trying to work out how old AI in the mainstream is now.
But if you think about the last couple of years, we're really thinking about let's just try and get AI usage.
Let's try and get AI adoption as almost like get it in the hands of developers, play around with it and see what works and what doesn't work.
It feels like more and more now going forward, we're thinking about how can we make it work effectively for us and how do we actually get the entire organization using best practices?
So let's flip it a little bit and say, what are, as organizations start rolling out AI seriously in terms of...
as a practice across every single development team, what would you say are the greatest bottlenecks today, which really, maybe it's platform teams, maybe it's other teams, that they're struggling with as part of that scale and that AI enablement across other teams?
The bottlenecks that...
become apparent are often deficiencies in just doing software development well.
And I think that we may probably discover in the next couple of years that we spend a lot of our time repeating ourselves of the things we've been saying for the last 10, 15 years of, is your CICD any good?
Do you have tests?
Do you have coding standards that people actually agree on?
So all of those things kind of get exposed.
And the DORA report showed that if you are not very...
mature in your software development practices, then agentic coding is likely to make you go slower.
If you're doing well in development maturity, then you'll go faster.
The problem is we don't know where that tipping point is.
So there are all these things inside the development process.
But even with successful agentic coding adoption, I've seen in a couple of customers and people that I've spoken to on the podcast, Probably shouldn't plug our own podcast there.
It'll be in the comments.
You're on the Aya Native Dev.
I assumed you were talking about that one.
Of course, you do have your own podcast as well.
Indeed.
But we've seen people there of you suddenly speed up the act of creating software and then products are caught flat-footed.
And they're like, crikey, we didn't have enough stuff in the backlog.
And that's hard to speed up because it's strategic and you need deep thinking and you need to be talking to customers, understanding their needs.
So we've definitely seen cases where software development has started to go much more quickly.
And then...
In people that do QA as an after-the-fact thing, that's been a bottleneck.
I'm not sure I'd recommend that pattern generally.
But yeah, product tends to become a bottleneck.
And then we need to empower those people to have more free time and to do the strategic thinking and the discovery work that presumably they really enjoy and they would prefer to be doing rather than writing JIRA tickets that developers aren't going to complain about.
How about you, Patrick?
I think if you look at the door, first thing is the adoption and kind of is it like cranking up new code that was almost like the developer productivity lines of code for a long time.
Then we learned, okay, it needs to work in production.
So we brought in the metric actually, like how many times do we need to rework it if it isn't good?
Like that kind of is your defect rate almost like that.
You kind of need to be on control both in generating and kind of making sure the defects are in balance there as well.
And I think what I like as an almost like new proxy metric that when you look at not just the users coding with AI, but if the agents are doing the coding, you look at the things that like how many, how many almost like turns does an agent need to do to do their job effectively.
and you can get that number down by, you can select another model, you can give it the right tools, and you can provide it better context.
So in that way, it is not anymore like, can I do better code in production, but can my agents do better code towards production?
So it becomes like delegated.
So you need to think out.
how effective they are.
And that's kind of a flip that I see now slowly moving from the dev working with AI to the dev instructing the agents to use AI to do it.
And that's kind of like an interesting proxy metric, I think, to track.
And I think there's something both amusing and mildly depressing about this in that when we have software factories and that idea matures and becomes more commonplace, we're going to have this infinitely tunable and measurable way of...
developing software and we see exactly how effective it is in that way that you mentioned of looking at how many turns are used, how many tokens are used.
And unlike with human teams, we're going to be able to A-B test.
We could have prompts and software factory setups that use one approach versus another and we can reset their memories and try it a second time and see, did that work better?
Did it not?
And you can't really do that with software engineers and real teams because you can't do a men in black and memory wipe them.
But the thing that's amusing and depressing is that now that it's agents doing the work and tokens, and the cost of those, everyone's going to be really keen on making sure the software development lifecycle was as effective and as efficient as it possibly could be.
But when it was humans, it was like, oh, well, Timmy's just lazy, and that's why stuff isn't getting delivered.
And they didn't focus enough on process and value stream mapping and all those important things to tune it.
It was just, oh, the humans can suffer.
But now when it's agents, it's like, well, that's going to cost us money, and we can easily see that it's costing us money, so now all of a sudden we care.
Yeah.
It's funny we face this challenge today because we have agents working on 160 plus organizations, which are widely different code bases.
And we see the ones that are better quality, the agent does a better job faster.
The ones that are lower quality, we have an onboarding flow that the agents get better.
After a few tasks, they onboard themselves to your repo.
And we see that in the difficult projects, it takes them longer to onboard to good quality.
But we focus on onboarding quickly to an organization.
We go top-down, so that means the business leaders we talk to, we cannot say, hey, do a sprint to refactor so the agents can have a better time, so everything will go smoother because they don't really care about that.
They just care about speed, like you mentioned.
So for us, it's been all about how do you mix in doing the user-facing work quickly while riding the natural code churn that the organization has to slowly refactor.
We have what we call the PAP framework.
And as you do work, you slowly refactor bit by bit so it's easier for the next agent that comes along in that part of the code.
But you do not come and say, hey, dev team, you have this task.
Make it easy for our agents to work there.
You have the agents doing the work in bite-size, going back to the PR fatigue in a way that doesn't inflate the PRs because then people won't go over them.
So it's a really, it's a balancing game.
Getting the leadership happy because features are being delivered.
Getting the developers happy because PRs are not too big.
And getting the, the non-technical PMs, the designers happy because they can do the work all while slowly pushing your agenda better and better coding standards.
So everything needs to happen kind of weaved together.
And I'd love to kind of, while talking about it, previously you mentioned about the kind of like...
the technical devs, the traditional devs, and the non-technicals, so maybe it's a PM or something like that.
When we talk about rolling out an enablement here, do you see different, tend to see different bottlenecks between a technical dev and a non-technical dev when adopting like a more of an agentic process there?
Yes.
Okay, so we'll go with the technical side.
The technical side, there's an ego that we all know.
We are working through that, whether it's with clients or whether it's internally, like the devs.
So for us, they interface with us.
They can use our platform.
Not to show off, but the PRs that come out of them versus out of the platform versus developers, often our platforms, they do a better job because they map your coding practices much more in-depth than developers, and they keep much better tracks of existing components.
We have clients with 2,700 components.
No developer will know which component to use.
We do know.
We do data science on it, and we can actually re-rank and pick between them in a much more intelligent way, but you have to come over that.
that ego of the developer viewing the PR.
So we tuned it.
We're a little gentle when it comes to them.
They're very fragile, aren't they, developers?
I'm not going to say that.
Did Patrick say that?
I can't remember.
And then on the PM side, it's a confidence thing.
It's like, hey, I'm okay opening a PR.
It's getting the dev team to play along.
We have some experiences where...
PMs open PRs and the dev team push back.
And then we asked the team lead, look at the code, is it good?
It's like, oh, this is solid code.
Why did the developer push back?
Because it didn't come from a developer.
So when that's worked out, we basically work on a conference of the PMs saying, hey, you don't need to write a ticket, that's not the deliverable.
You don't need to make an HTML prototype that is disconnected, that is not the deliverable, that's effectively a ticket and a Figma.
You need to actually prototype inside the code base.
You need to edit existing stuff in the code base.
hey, it's possible, hey, it's feasible, hey, effectively it's really easy.
And when you work with the agent, it already knows the constraints from the code base built in when you interact with it.
And yes, when you click send to dev and it goes to the developers as a PR, be proud, no be scared.
And that's like a big enabler.
It's almost like an imposter syndrome, I guess, for non-technical folks to feel like maybe they're resistant to do things with confidence because they feel like they're not a dev and a dev's going to scrutinize it more almost.
They're scared.
They don't know how to answer technical questions.
And that's a part of the process we're going through.
You don't need to.
That's the beauty of the agents.
They know the code base.
They manage it for you.
You need to give us the product intent.
And that's what matters.
And the agent delivers something visual.
Then it's a part of their app that is rendering and running and they can interact with it.
If you do the product review or the design review, depending on your persona, you should be confident.
The code works the way you wanted it to.
The code quality is there.
You don't need to worry about it.
And hey, if there's like judgment calls on the back end side, the developers will take it.
So let's move in a different direction now where we're going to talk about ways in which agentic development or rather patterns of agentic development, how people are doing this in the real world and what we see from people externally and in what they're saying they're doing and how they're approaching it.
Patrick, you recently...
created a really interesting tool.
It's called the Tesla Patterns site, tesla.io forward slash patterns.
And it's a ton of research in the background.
I think he uses the Carpathi wiki approach in the background.
And it provides a bunch of patterns that people can look at and say, okay, I can see that people are thinking and doing things in this way, this approach.
And, you know, sites, various social, you know, I guess, ways in which people are going into more depth to say, yes, I did it like this in this infrastructure in this way.
Talk us through a little bit about that and maybe highlight a couple of the more interesting patterns that you found.
Yeah.
I think one of the reasons is that when you do surveys right now, you get like an enormous spread.
Like, you know, person A is savvy, the other ones are not savvy, and then the effectiveness, and yes and no, that's hot.
So I figured that what if we...
look at the socials on how they do things and kind of capture that signal.
Now, that has a little bit of a problem.
It could be vendors saying how they do great things and practitioners.
But kind of the old idea is that we filter out, like how some of the people are using this.
And, you know, if there's many people doing similar things, that boils down to kind of a pattern.
Now, one of the surprising and unsurprising things is almost like what is good for AI is actually good for a human and the other way around as well.
Like, given that you're working with an agent, if you actually give your agent good documentation, lo and behold, it performs better.
If you give them tests, then you actually know whether that's still working, yes or no.
So if you have observability, it could expect itself.
So that's kind of like, you know, when you start thinking about like these analogies on like what is actually good and what is not.
You can go up to rituals, for example, on the team.
We have a team lead and they're used to doing scrum and retros and kind of those things.
But what if you would have an agent review kind of at the end of the sprint?
What were they missing?
How effective were they are?
But it's not also what do you do as a team lead?
You set a goal.
for a person.
So the goal sets the measurement.
So that is another account.
So you find like people doing like those little ceremonies or for example they do writing almost like context together as a pair so you get nuances.
So that's the pair programming but now with context.
So that's kind of interesting that those signals are bubbling up.
in a certain way.
Which is good because it actually goes back to what you said earlier about it's good software development practices.
Everything that you mentioned there should be things that we are doing today manually, but it's about almost like, can I say the word, agentifying, agentification?
You know, essentially making an agent take care of all of that.
Yeah, yeah.
But I think it's, and you already mentioned that, if you want, yes, it gets used.
gets a little bit getting used to as you know what's the new world as a PM and an engineer and kind of I don't I feel okay contributing and working in that but for example if there's one advice that I would give a developer right now moving to agentic it's just the most important thing would be don't repeat yourself like don't keep telling the agent what to do but write something down if you're using a tool to verify with its dueling give it the tool so you're not having.
So kind of that is a very strong, like almost like mindset.
And it's different from, oh, let's collaborate, like not collaborate, like just have it do all the work.
And that pressures you in all the correct engineering points.
And I think when AI and coding came up, oh, you know, the typical is like, this is the end of coding and there's no more engineering.
And lo and behold, you know, if you really want to get all the value of that, It is good engineering that we bring in there.
And the engineering becomes creating the software factory or the assembly line or whatever you want to call it.
I can imagine that for a lot of, like I can see a bifurcation for developers of the more product-minded ones who like achieving outcomes and getting motivated by, I delivered a feature to users they find valuable.
They're going to kind of become more product-focused.
And the ones who are like, I love making nice, tidy, neat code.
they're going to be more motivated by, I'm going to make the best agentic software process that I can.
I'm going to tweak the dials and the knobs on how the software gets generated.
So that kind of engineering mindset doesn't go away.
It's just you're not applying it to the code.
You're applying it to the machine that makes the code.
Exactly.
Yeah, yeah.
And I think that is interesting about, we had that when we went to the cloud, everybody wants to rebuild their own kernel.
Why?
It's good enough.
We don't need everybody to build their own kernel.
So if there was one advice that people come up to me over all the years of DevOps and they say, but yeah, you know what?
We're special.
It doesn't work for us, blah, blah, blah.
Guess what?
The industry kind of said that loop is actually universal.
It can work everywhere within the software development.
And so my advice then to platform team would be saying, okay, your number one focus is telling people that they're not so special in what they're building for code, but it's their mindset that is the difference in kind of how they approach the problems and what they bring into that mix.
And it's really weird because everybody's repeating this, which is great.
Like every new technology has a learning phase and we all have to go through the motion and kind of say, hey, oh yeah, now I understand.
And yes, all the stories about like, I vibe coded this app, like all great, but how do we bring this into the maturity of an organization?
And that's only by saying like, while we probably don't need to, every developer needs to build their own pipeline, that's probably not very effective.
And, oh, it's actually the same thing over and over again with two variables or something.
And that's kind of the mindset of the platform, the reusability across different teams that goes in.
I can share a little bit about our process internally.
And I guess I think the R&D leadership perspective I try to push.
The way I think about it is I go back to pre-gen AI, and I imagine I had all the manpower and all the time in the world.
Somebody opens a PR.
I would want to map out, is it high risk, is it low risk?
PRs merged.
I'd wait two weeks.
I want to go over the logs and see, did what they do, was it successful, was it not?
Is there some edge case they didn't cover?
And basically with that mindset, we say, okay, now you have Gen.AI.
You have agents.
There is no excuse not to do everything you ever thought you'd want somebody to do and then automate it so it happens by CI automatically.
So, for example, here, all of our PRs get labeled by risk automatically by an agent.
Cool.
Then they go to staging.
They live in staging during the QA process.
Before we release, we have another agent that goes PR PR, sorted by the risk level, and goes through our logs.
Our logs are queryable for agents.
And they actually go and check every single PR.
Hey, what was the intent?
Hey, what did the logs look like?
Did it cover what it was supposed to change or fix?
Was there some edge case we missed?
And it's been amazing that it's caught more and more stuff.
Like the high-risk PRs often do miss something, even though our engineers are...
agentic and how they work and everything and we have QA and we have testing and we put in all this effort but hey when you retro it a week after it's in staging you find more stuff and often it's like easy stuff like the agent knocks it out it's like oh this edge case pops up once and never because we're not a terministic product because gen AI and LLMs but hey now we see that and let me capture it too so that sort of mindset of if I had all the manpower in the world what would I tell them to do?
Okay, we'll start having agents do it once, twice, automate it for us, we put it into our release cadence, and suddenly you just get better and better quality, and you move stuff that you would need to think about, or your developers would need to think of, just move it.
It's just automatic.
I love that you bring up the risk thing, because there's the belief that the digital factories will churn automatically and all the things, but because when you know what the risk is, again, translating it to a management world, Different risk levels require different management techniques.
Like, are you micromanaging?
Yes, probably the risk is very high, and then you need to do that.
If you have guardrails or something that is not that important, or you have some way of mitigating this in the product, by all means, then we'll get the feedback later.
So that is the levels, and you can only do that if you start thinking about risk levels of things pushing through.
We even couple it to our PR fatigue.
So high risk, get more human resources to go over it.
And we even couple it to our token budget.
So low risk, we're not going to spend a lot of effort to check it from an agentic LM spend.
So just a way to prioritize everything.
Sounds a lot like running a business now, right?
When you talk about adding this into, like, you know, instrumenting this through your CI, is CI still fit for purpose in this flow?
Or are there alternatives or what changes we need to make to our CI style process to make it better for agentic workflows?
Oh, there's like key transformations you have to do.
So one key transformation is whatever your observability platform is, you need to make sure it's easily queryable for an agent so it can verify both while it's doing a task and also retroactively on entire, I don't know, dev environment, staging environment, prod environment.
You need to make it really easy for the agent.
And we have skills that also tune themselves so we continuously get better, especially as our product evolves and our logs change and the way you query change.
So that's like a key part.
RCI is very agentic heavy.
Even the way you push and open a PR is extremely agentic heavy.
And all the actions that we put in from a DevOps perspective are also really agentic heavy.
That's just a way not to have, like you want it to end, you want it to test, but you want all those mechanisms for feedback to the agent.
So the agent works in higher confidence.
But you've got to think of how is the agent stepping in and how are they doing more than just writing code?
How are they your confidence layer that there is quality, The risk is managed.
The stuff is moving faster.
We even have agents help us prioritize.
You have 20 PRs open.
Which one is high importance?
The agents will tell us and push the reviewer to act on it.
Do you see it almost like this series of actions that need to happen?
Or do you feel like it's more fluid than that where you have different agents providing you with different data to make decisions more integrated, more collaboratively?
So I'm going to go with your answer, which is it's fluid until I figure it out.
Yeah.
And then it's hard-goated because I don't want to repeat it.
I want to get that mental load off.
I've got a strong hunch or strong opinions really on this in that...
When you're a software developer and you're trying to work out, like you're implementing a feature, like one, just make it work, then make it work maintainably, then make it work readably, performantly, securely.
You've got all these different lenses through which you have to look at a code change.
I think that what we need with agentic development is to apply all of those lenses deterministically.
Instead of hoping that an agent is going to figure out all of this in one pass by sticking all the things in agents MD, instead we can have some kind of CI-like process.
I've built a little tool called Assembly Line that does this, where you make a change with your agent or the agent makes a change by itself, and then another agent gets spun up with one deterministically defined prompt of check for dead code, check for missing test coverage.
Check for this, check for that, check for the other.
And to your point, you've got infinite engineering resource.
So why wouldn't you look at a change through all of those lenses and then try and provide feedback as fast as possible to the agent that was doing the main kind of set of changes so it doesn't deviate too far and can fold that back in?
So I think we need to, as an industry, as practitioners...
figuring out that the flexibility and power of LLMs to work in a non-deterministic fashion is really powerful, but we don't want that all the time.
I know of some people that are using agent skills to figure out how to deploy stuff into prod, and I'm like, just build a platform.
We've had platforms for 10 years.
They're nice and deterministic and straightforward.
Maybe use agents to build your platform, but don't replace your platform with agents burning tokens, reinventing the wheel.
There's a place for the determinism.
And I think it is in that kind of figuring out, like you explore and then you figure out, okay, right, that was the right place where we should have run that kind of check or a pass, a sweep over the code base through this lens.
I have a ton to say.
Okay, so excuse me.
So we have two harnesses in my organization.
We have the harness in the product and the harness for the R&D team internally.
So it's beautiful because what you said is stuff we have in both harnesses, meaning we do not want the agent to go through this route to exactly what we want.
We want the agent to do a big step this way and then...
that step makes sense for it statistically.
And then a big step this way, again, it makes much more sense to it statistically, you see that you're no longer fighting with the LLM.
It does not want to do one, two, three steps in one.
It wants to do one, wants to do two, wants to do three.
You just get better quality that way.
And it's funny, you call it assembly line, we call it nudging.
But in that sense, you get a very clear workflow and you no longer fight the agent and how it wants to work.
What we do internally for our...
R&D team, the R&D harness on the product.
Sadly, we have to factor in the human element just because we do.
So we developed...
It's like the slowest part of the whole machine.
So we developed kind of pump is what I call it, which is plan merge polish.
The code is moving so fast, you can't have an open PR.
Meaning if it's an open PR and I want to do QA on it and I want to do like a product review and I want to do a design review, that is all lovely.
By the time those are done, the code is going to have so much conflict, everything looks different, I have to restart.
I am going to be so, so happy if we manage to go back to trunk-based development as an industry.
Like pull request-based workflows inside enterprises are a damn silly idea.
They make great sense in open source repositories with people who are not strategically aligned and you don't trust and, you know, it makes sense there.
But like that...
Because of that speed change that you mentioned, you've got to get the feedback to the agents working on the code as quickly as possible.
And that kind of goes back to a point that I think is running through all of this, that the fundamentals of what made good software delivery, what made good methodology, those haven't changed.
Things like fast feedback are important.
They always will be.
Having feedback, being able to tell, the point about observability, that the agent can tell whether it's broken something in production.
Having all those feedback signals.
was important before for humans and continues to be important for agents, if not being even more important because of the rate of change.
I'll expand.
100% agree.
The issue there is also that as soon as PR is merged, the landscape changed, the code changed.
So the next agent that starts will do a different job.
Everything impacts the statistics of the path that LM goes through.
So for us, we called it the M in the pump, the merge.
As soon as the PR is open, we want to get it merged as fast as possible.
So hey, human reviewer, look at it right away.
Get it in.
We need it to pass our CI.
We do not need it to pass QA.
We do not need to pass the product or the designer.
Everything is feature flagged.
Then we go into the last P, which is polish.
Stuff is in the code.
Every new feature that gets developed is off of the same code.
And our product manager can go and review it, and they can polish it using our platform and get the UI and UX exactly right and see all the bad decisions from the user-facing perspective that the developer made.
Well, the developer got the feature working.
It did not get the feature polished.
But that's fine.
We have other personas for it.
And now they can do their work in a way that is smooth within the evolving code base.
It's evolving really quickly.
So that kind of feature flagging, merging quickly with feature flags enables us to do it kind of two-step.
There's no one PR.
A feature should take three or four PRs.
One PR from the developer.
One PR from the product manager that changes functionality.
One PR from the...
designer that tunes the UI weeks and then maybe one last PR from the QA that saw some sort of edge case.
And in terms of the human reviewer being the slowdown there, I had a really interesting chat with Ryan Lippopoli, who's a member of technical staff at OpenAI.
And one of the things that they're doing in some of their projects is to entirely remove a human reviewer.
How much does that make you go, oh, gosh, that's horrible?
Or how much does that make you think, oh, wow, there's a real opportunity here in certain cases for us to actually allow either an AI or tooling to perform those reviews and give us that feedback, maybe for different levels of criticality of change?
But is that something that you look at much?
Oh, all the time.
So again, it goes back to the risk conversation.
High risk, no way.
I want a human person.
And I also define like a risk is also by what part of the code base they're touching, how sensitive it is.
But beyond it, I'm trying to automate the PR commenting and comment resolution process, which means I have agents that have modeled every single developer on my team and their comments for the last few hundred PRs.
And I know each and every persona.
So as soon as a PR is open, we have a whole prepared PR, which is a bunch of agents that check it in all the ways that we've modeled it.
We need to check it.
Once a PR is open, we launch different agents that model different people from our team and comment as if they were those people and does a surprisingly good job.
Then we trigger agents again to resolve those comments.
Then we trigger the retro to see are there any error logs, did stuff actually behave like it wanted to.
And again, if the risk is not high from the PR's nature, I'm moving towards being able to skip the human reviewer at all.
Again, I have developers on my team.
I have some developers on my team that are absolutely against this because they want the developer to own.
So I'm going through the motion, massaging them and getting them comfortable with the idea.
But I'm definitely pushing for that.
And it depends how reversible it is as well in terms of...
It's the guardrails, right?
And it was the same thing like when everything was being automated during the DevOps transformation, people said like, but if I did change one thing, I can delete everything now, automate it, right?
But it was the harness of tests, the CI, CD, that kind of was like the fail system of...
And then the narrative was you can have a junior come in.
push to your, you know, kind of your main and like your harness, your test harness will catch it.
Like if that's a good one.
And but what you saw is that there was still a difference between I'm going to do changes in the code of the compute to I'm going to do changes on the data scheme model because people felt the risk was higher.
So you can see kind of similar patterns on like the choices that you make.
But it's a mindset, like how far do you go?
There's also costs to kind of doing the automation.
And you mentioned a lot of things like, yeah, I have an agent for this, for that.
You know, putting everything in place is currently like a lot of work, right?
And yes, you can buy your product or your product or like everybody's product.
But that kind of is, you know, funnily enough, we hope everything will be...
more simple 5Q and everything will work.
And what we end up, and that's usual for any technology maturing, is that when it's matured, it's even more complex when you want to reach the next level of kind of full kind of, you know, power of the new technology.
One of our customers at Aervo, they're doing some really great work where they've got a team that has gone philogentic.
They've changed their development process as to...
embrace ai in many different ways requirements gathering transcription story writing implementation all of those kind of things and they are now going 72 times faster so they have done eight years worth of work and they're on track to complete that in 11 months with a third of the people so they've done great things there they've given up on human code review but each individual developer they take an epic and then they some of them use beam ads some of them use spec kit some of them vibe Before they self-merge their PR, they run about seven or eight code reviews, different tools looking at different things.
And their confidence is, they've been doing that.
They're now confident enough that they don't need humans to look at it.
And I think there's something interesting.
I don't know how true this is.
I mean, data is definitely like when you screw up data, your data is gone.
Do we need to move to event sourcing and having a replayable history of all data transformations?
That's one thing.
But like generally, As a trend, we've always seen that trying to prevent bad things happen by putting a check in place is the wrong answer.
Like, with data and databases, We went from having transactions to prevent inconsistencies happening.
That was slow.
It doesn't scale.
So what was the answer?
Eventual consistency.
And we'll reconcile after the fact.
With the cloud native transformation, we found out that you can't have a thousand microservices and then integration test them all.
You just have to deploy them and find out what happens.
So you need great observability and you need to be able to fix forward very quickly.
So again, you let the bad thing happen, but you react to it more quickly.
I imagine...
that where we might end up with agentic coding is something like outcome-driven development where you've got observability of the business impact you wanted a feature to have.
Like, is the application doing the things the application is supposed to do?
Are users able to do what they should be able to do?
And then we have software factories chucking features out, and then you look for regressions in, like, okay, the number of checkouts on the shopping cart has dropped.
reacting and fixing forward at that point.
Is that something you feel users are going to accept?
It would be interesting to see in that Do we all end up getting used to buttons moving locations and software being less kind of...
Who was it that coined the term?
I think it might have been Steve Yegg.
Might be misattributing that.
But of non-deterministic idempotence of this idea of you have the same spec, same product requirements, you run it through an agent twice and you get two slightly different things.
but you enable the same user behavior.
Just as we've got used to waking up and finding out that Apple have decided to change how the photo album works and our elderly parents are now very annoyed, do we end up with something like that on a more daily basis, like this rapid churn of features changing and being slightly different, but I can still achieve the outcome?
I don't know.
There was a testing Facebook.
for example, like it has so many rules, so many features, writing that kind of in a consistent test.
And the way that they did it was they added agents that were using Facebook and they had their own rules.
Like for example, like that agent should never get a friend, like whatever, like it should not be listed visible in anything.
So they had these kind of like flags out and by all these kind of things.
But they couldn't only make it work when they hooked them in on the production system.
And that was a novel way of thinking of, instead of the rigid test flow, almost like you track whatever behavior is going, if there's bad behavior.
It's a different way of looking at observability instead of your logs.
But kind of that was...
An interesting way.
Like a digital twin is also something I think will have a more future is kind of predicting things, what is going to happen, and then the risk comes back like, okay, I have a model.
I can run it through.
I can see what's happening and kind of those things.
But, yeah.
It's funny.
I thought about Facebook too, but from a different perspective.
They have a huge infra setup for A-B testing at scale.
Any developer can do whatever they want, and it'll go and get tested.
A-B testing facilitated again on scale.
And I think that's kind of what you were going for.
How can you do it on scale and just see what wins?
Which sounds great on paper.
A lot of the organizations we talk to don't have the setup for it.
And then we get some organizations that come to us.
Hundreds of organizations that are, I think the one I talked to last week was like, just shy of a thousand people.
And they said that their product is just downward spiral because everybody's agentic development is moving so fast.
There is so much PR fatigue.
the stakeholders, like the PMs and the designers, stuff doesn't go through them anymore because everything's moving so fast.
And then on the call was the head of the front and in front.
And he says, yeah, I'm on calls.
And I discover new features on the call that don't look anything like our other features, that don't even match our design system, that functionality-wise don't make sense to me.
And I find them out with the client live.
And then I ask the developer after the call, what is this?
He's like, oh, this got merged two weeks ago.
And then I asked, did anybody look at it?
Was there a product review?
If there was like A-B testing, then at least it'd be guarded in some sense and not just immediately impacting all of their users.
And again, huge organizations that's been around for over a decade.
But they have adopted agentic coding really quickly without amazing guardrails.
And now they're kind of spiraling back.
They're like, wait, wait, we got to control it.
We have a huge brand.
We have users that depend on us for years.
I can't move the button and make it hidden because this is a core workflow they need to work through quickly as a part of their business needs.
But it's almost like it's the, you know, is the UI really going to be the way we, you know, interact with software these days?
An agent will be absolutely fine with it.
It's us, again, pesky humans, which is the problem.
And how much of this will actually change?
And you actually think, no, I don't need a UI on screen.
I need to ask certain things or I just need my agent to go ahead and interact in a specific way.
So I don't know.
I wonder if in years buttons and things like that will just be replaced with completely different ways of us interacting with apps through...
But we'll still be a visual-oriented creature as well, right?
We can't just say we're just typing in the words.
There was a talk here as well from TL Draw, kind of interesting that, you know, having your agent giving it a whiteboard, because sometimes you can express things easily in the whiteboard and kind of like typing.
And that kind of fluidity is also, I think, in the UX.
Like sometimes it is going to be, I just want to ask a question, you reply back to me.
Sometimes just show me, like explain it to me.
But it might be interactive and kind of going through that.
I think that's also the pendulum, which was funny, like IDE, CLI, and now we're back at Orchestrator Review.
tools because we need something like we need to see what's happening and we can't just like do that but the agent might crack along in their own coming communication and find the best way of kind of dealing internally but as long as we need to make call shots on kind of those kind of risk we need to be informed and somebody explains it to me like what are we doing here and that's a challenge if we don't do it ourselves anymore so we've talked a lot about good patterns about ways in which people should change the way they develop software and to essentially directing people and groups as to how they should change.
But what is it that organizations, enterprises are doing today as a part of their migration, part of their adoption that is actually just a mistake?
It's wrong.
It's been proven that dragons live here.
What should they not do that they're doing today?
I think I'd give them the advice is that there's still the tendency of, if we're doing a constant conversation with our agents, we're in control.
I think there needs to be, for them to get into the next mile, they have to let go.
But it's almost like you have to mentor the other one instead of doing the work.
And that's kind of the mental...
challenge that I think a lot have right now.
Stop doing the work yourself.
And it's really weird if that's what you've been doing for so many years.
And what a lot of developers actually hold important to themselves because it's what they feel sometimes actually makes them the expert, right?
Yeah, correct.
It's interesting.
I have a few different answers to this.
The first one is you want to control your AI spend.
People are releasing crazy budgets and everybody's trying to token max every single developer.
And if you, I think the right way is give the developers the tasks that interest them, you will see they're way more engaged with the agent in the session where they're building it versus give them the task that the developer is not interested.
And then they will off-source all that mental load to the agent.
One, you get poor quality PRs.
But two, you get way bigger token spend.
Like imagine that something's out, the disinterested developer will just say agent test it instead of testing themselves and seeing exactly what's wrong, saying, oh, this is the issue, let's fix it.
They will let the agent do seven sessions of QA and the token count for the session goes crazy because the developer is not interested.
And then, so that's one part.
Like, yes, people can move faster.
Giving them tasks they don't care about is a great way to spend your AI budget really quickly.
That's one.
The other one that we've seen a lot of organizations do is they want to jump in headfirst to AI native.
And they do that by giving cloud code to everybody.
And the cloud code has code in the name.
It is great for developers.
I love it.
It is most definitely not good for non-developers.
And if you give them that, you will get PRs.
One, they don't know what they're doing.
They don't know how to set up an environment and they're going to drive your dev team crazy asking questions, even if they feel comfortable enough asking questions or they'd just be not doing any meaningful work.
But if they do open PRs, we see garbage in, garbage out.
PRs that the VP, we're on calls with CTO and VP product.
And the VP product...
shows off to us saying, hey, I opened a PR yesterday.
And then the CTO sitting next to him says, yes, it was crap.
And that's also a great way to not go about it.
So when you do enable them, you want to get the tool that is really optimized for that user.
Again, we optimize for PMs and non-developers.
But giving the people the tool that's not meant for them is not the answer.
Again, you use your AI budget really quickly.
And you get frustrated people all around.
The PR fatigue is there on the developer side, and PMs that open really poor PRs are increasing the PR fatigue.
They're getting more frustrated themselves, and you get uber scenarios where everybody's just disappointed.
I think in terms of what people should not do or stop doing, like the whole piecemeal approach to last year, 2025, was very much the year of CTOs going, well, I let people use their own tools and whatever they're comfortable with.
that is going to end up being inconsistent and you're not going to get the organizational learning that you need.
So a firm mandate that like the AI bus is leaving, please get on board.
That's been shown in research by multitudes, I think, to be a strong indicator of success of agentic coding adoption.
The counterpoint to the token maxing is also don't go to the other extreme and set a really low limit.
I've seen places where there's a 100 euros a month token limit and like 100 euros you can't really do very much.
So it ends up adding cognitive load of people like, oh, should I use CaaS for this?
Or should I, is this the most important thing I'm going to do this month?
So having something in between would also be sensible.
And then maybe rethinking what a test is, but to invoke Beyonce, if you liked it, then you should have put a test on it.
If you care about a thing, then you need to have something in place to assert that.
Whether that's functional tests, whether it's a sweep of an agent with a particular lens looking at a concern, whether it is connecting up via MCP or agent to observability so it can figure out the thing that it just...
built and got deployed is now breaking.
Making sure that there are those feedback loops in place.
If you just have an agent on its own and you expect it to get everything right first time without being able to even perceive the mistakes it might make, then you're going to get bad results.
So making sure that you think about what you value in your code, that it's architecturally sound and the code quality is high and the test coverage is good.
You need to work out what you value and then make sure that there is something in place to prove that that is there.
It's not quite testing, but it's like thinking about the non-functional aspects of your code as if they were tests, if that makes sense.
Can I expand slightly?
I love that point.
I think something that's really important is also the temporal mindset and the context-hinting mindset.
which means every test needs to stand the test of time, needs to be evergreen.
But you also need to think of it as an opportunity for a context hint, meaning when something fails, the test has enough context in whatever the assertion is, the comment, the error response, whatever it may be, so that the agent that gets it will know how to act.
So lazy developers just say, cool, I covered it with tests.
No, you need to cover it with tests and you need the test to have the context hints that the future agent that will see them will know what to do with it.
That's been really impactful, getting the agents to actually manage in sophisticated code bases.
Yeah, and it's exactly the same as we were all saying earlier of like, that's always been good development practice.
Like, there's nothing worse than as a developer, you run the tests and you just get like a SERP true failed and you're like, great.
Or 1317, yeah, what?
Yeah, great.
What do I do with this now?
So yeah, it's important stuff.
And thank you for quoting Beyonce because I've been trying to work out how I increase that in this podcast, but finally someone has.
So thank you very much.
Wow.
We've talked about a lot and it's been super, super valuable to have you all kind of like discussing on these points.
So I appreciate all your input on that.
Thank you very much for joining us at AI Native DevCon.
Why don't you very briefly talk a little bit about some of the sessions that you're each giving.
Patrick, why don't you start?
Yeah, I'm talking about the different layers from up to a dev to a team to a platform team and a VP.
What's the mindset that you need to have to go into this new agentic world?
I'm not talking about generic transformation, like do a hackathon and so on, but what is really like, how do you embrace the new thing?
So that's my talk.
Awesome.
So my talk is about what it's like to become more of an AI native organization where your PMs, your non-developers are actually opening PRs.
and how really the metric you should be actually should be tracking is the merge rate.
So the PRs that they open, don't say how many PRs did you open.
Look at what the merge rate is.
That's how you will know, are they opening quality stuff or are they just increasing PR fatigue?
Yeah.
I was co-presenting with Tamarj Mai from Adervo, and we were talking about the story of how we upskilled 120 of their developers in agentic coding.
And the most important thing about that talk is you get to see me in a day-glow fluorescent tracksuit.
So people should definitely check that out.
Is there a Beyoncé quote in that one?
I don't think there is.
Shakira, maybe?
Sadly not.
All right, all right.
So why don't we, well, all of these talks are recorded, so you're very, very welcome to go, and we'll add some links there in the show notes.
If I or our listeners can only go and watch one talk, which one would you agree that they should?
No, no, let's not do that.
All of them, all of them.
Thank you very, very much.
Really appreciate you being here and sharing your time with us.
Thank you very much.
My pleasure.
Thank you.
I hope you enjoyed that.
That was really great from my point of view.
Would love to hear any feedback on that.
Please let us know.
podcast at tesla.io and tune into the next episode.
Bye for now.
Hope you had as much fun listening to this episode as I had recording it.
And if you want to hear more from the guests featured in this episode, we have their full talks linked in the description below.
So that's it for this episode.
See you next time.
The AI Native Dev is brought to you by Tesla, the package manager for skills and context.
Your hosts are Guy Pajani and me, Simon Maple.
Our producer is Tom Dowler.
The AI Native Dev is not just a podcast, it's a community, and we host monthly meetups at the TESL offices in central London.
Visit tesl.io forward slash community to learn more, and I hope to see you there.
