# Datadog Lessons For Scaling Agentic Coding

**Podcast:** The AI Native Dev - from Copilot today to AI Native Software Development tomorrow
**Published:** 2026-08-04

## Transcript

The hunch we've had is like maybe actually the context that was written like a year ago, you know, when it was like cursor, Sonnet 3.5, it's still relevant today.
And so we started to look into it and then we're like, well, what if we delete the whole context from the repo?
Like what would happen?
And what happened for us is actually like the evals started to perform like much better, which was counterintuitive, but it's just like, yeah, that file was pretty big.
and was listing stuff that was not like necessarily useful.
And so that was really interesting to see.
The AI Native Dev is a podcast for developers and engineering leads at the cutting edge of AI and agentic coding.
Join your hosts, Guy Pajani, and me, Simon Maple, every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.
This is the AI Native Dev.
We just wrapped up two amazing days at AI DevCon in London.
But the great thing is that we get to do it all over again in New York City this November.
You're absolutely right.
We're going to be back in the city that never sleeps on November 3rd and 4th for more amazing sessions, really engaging hands-on workshops and much more.
Yep, all that great networking, partying, eating and drinking that you've come to expect from AI DevCon.
We think we have one of the best hallway tracks in the business and it's the perfect complement to our incredible speakers and presenters.
We'll both be in-person and virtual with live-streamed access to all main stage keynotes and talks.
Sign up right now for our Super Blind Bird ticket for just $100, only available for a limited time.
We're really excited to be headed back to the Big Apple.
We hope to see you all there.
Hello, everyone.
Welcome back to the AI Native Dev podcast.
Today, we have a really interesting story about how to build out agent enablement and learn from it, looking at Datadog as an organization that really dove in and invested in agent enablement and has a bunch of learnings to tell from the journey.
And to guide us through all of that is Simon Boudrien.
I hope I'm not sort of misread that too badly.
Simon, thanks for coming on to the show.
My pleasure.
So, Simon, we're going to...
do this justice and sort of go in chronological order on it, but just to sort of give people a little bit of a teaser of some of the highlights, I know, as we talked through the journey that I really liked, we'll talk a little bit about context, but actually a fair bit about context, you know, where did it work for you when you created it, where did it create kind of context rot and exaggeration, and, you know, I don't know if it's dramatic or standard case, but a learning about cases where you...
Actually, an eval showed you that if you deleted all the contacts, it actually did better than before.
So that'll be great.
We'll talk a little bit about org and how do you structure and who owns what, the importance of evals.
We'll talk about where the bottleneck is, what are the problems really that are worth solving, because there's infinite work in the world of already doing it.
So I'm keen to share all of these stories.
For starters, can you just tell us a little bit about what is it that you do, the scope of responsibilities, and maybe a couple of words about Datadog and the large group of engineers that you're trying to optimize their work for here?
Yeah, for sure.
So I'm a director at Datadog.
I manage a part of the SDLC org.
We're called for the Language Foundations.
And so we are serving about 4,000 engineers here at Datadog.
And the way that I've described our scope is basically kind of the front end of the SDLC.
It's like everything people interact with day-to-day and how they do their work.
A lot of that falls within my group.
And a year ago now, we launched a new AI DevX group to focus on AI tooling basically here at DataDot.
Yeah, that little sort of a small initiative there of, you know, maybe this thing has legs on it.
Cool.
Okay.
And I like the definition of like the front end of the sort of the developer tools on it.
I guess when you have 4,000 developers, there's a depth to that as well.
Walk us through a little bit of the storyline of, you know, as you started to invest centrally or in general, like how did the kind of engineering organization and Datadog kind of approach or start opening up the use of agents and a little bit of that timeline of how it evolved to today?
Yeah.
Yeah.
So like for us, it really started in the beginning of like 2025 when, you know, Vibe coding as a term was like coined.
There was much more talk about it, about this being viable and starting to be something that really we could use AI to write code.
And so for us, the big driver there has been Cursor at the beginning.
And so at the beginning of 2025, we launched a POC internally with Cursor.
Initially, we were aiming to get 100, maybe 200 devs to use Cursor and just get some signals of how it worked for them.
And it just took off from there.
We had so many more than 200 people.
We quickly had 1,000 devs using Cursor almost day-to-day within a month of launching for this POC.
And so for us, it was a pretty big signal that, yeah, there's truly something there.
There was a lot of people that liked it, fully made value in it.
This is pure import, right?
You opened the door, you said you can use Cursor, and everybody came running it.
Correct, correct.
Everybody wanted to get in and try it out, basically.
And so that's where we started.
And with that pull towards more AI tool, we were like, well, yeah, this is going to be transformational to this industry and how we do our work as engineer.
And so at that point, we decided to, okay, we will probably need to create kind of a team.
to own those tools and own the experience around AI used for development.
And so at that time, we started to recruit a team that pretty much went through the whole summer.
Before the recruitment of the team, the sort of the we over here is basically you and the people looking at developer experience and front are the people, I guess, before it was designated AI.
Yeah, yeah, yeah, correct.
Correct.
It was like folks within my org rebranded for the POC.
It was like a staff engineer mail who took for the shepherdship of the cursor in POC.
Yeah, and I liked how you sort of told me that at some point this was a demand was so kind of off the list that you had a product manager that found themselves manually because like you weren't ready for the avalanche.
And so someone was like manually adding people's emails in to sort of add them to the cursor list of cloud users.
Yeah, yeah.
That was with like Cloud Code over the summer.
When like Cloud came out, like there was a second wave of adoption and like that tool.
So at first we had issue to set up for the SSO.
I'm not exactly sure why, but that didn't work.
And so we needed to add manual email to like Cloud Code basically to get access to people.
And yeah, PM took that role of like basically managing access to that tool.
over the summer until we could figure out the SSO bug we were running into.
Somewhat ironic in the era of AI.
So sorry, I kind of disrupted your flow a little bit.
I love that story too much.
So you had Cursor, you decided to hire someone, and then you went on to start Floodcode?
I guess, how did Floodcode enter the scene here?
Yeah, so because the product were already working with the Entropic models, we had connection already in this contract.
And so just like someday, clock came out and people could basically use it because we had api keys um and so and so yeah it just was available overnight and i think you know people pay enough attention to like the new cycle with like people are like on like saying there was a lot of hype around clock code at that time and so it basically organically grew from that point Would you say the Cloud Code adoption, this is history at this point versus present, but would you say Cloud Code adoption was as fast as the Cursor one?
Because Cursor, they really nailed the kind of low friction.
You're already in the IDE, we'll give you superpowers.
So why wouldn't you want Cursor?
Well, Cloud Code was more of a big change to it.
super compelling, but still the appeal was the same or was Cursor still adopted faster?
That is hard to say.
I'm not sure.
I'd say both at that point, towards the end of the summer of 2025, because I think the story changed later.
But at that point, they were pretty similar.
And I think they were catering to different personas.
Let's take people, especially for us, maybe piece of context, we have huge motor repos.
And so we know like VS code on those tends to perform pretty badly.
And so the base of cursor and some other mono repos was causing issue for adoption of just like being not performant enough for the large repos we're having.
And so there was like some of those people had a much better kind of experience with with clock code running in a terminal where like you don't need to have the IDE running.
And so for a lot of people, that was more compelling, I think, just from like a performance reason.
And so I think like that drove a lot of people there.
And then all of our power user quickly flipped over to Clawcode because I think what they were seeing is like they don't need the IDE open.
You know, they could do most of their work without an IDE open.
And so for them, that adoption for the power user was like pretty clear early on.
Got it.
Yeah.
And so this is, so we're about a year ago, we're sort of like approaching summer of 25.
Yeah.
At this point.
Okay.
So talk us through so cool.
So you've had that, you had a cursor just sort of lumped in and you decided to hire while you're hiring, you sort of have this sort of cloud code, you know, land in your lap and you're managing sort of added, you know, with that sort of SES, SSO.
And then you said you started to build a team.
Yeah.
Yeah.
Yeah.
We started to build a team.
So.
So like one of our early champion was in EM here at Datadog.
We offered him to do kind of a team switch and like build this new team.
We started to recruit people for that team.
We had a few trends for it come in as well.
Yeah.
And so, yeah, we started to put that team together.
We wanted to do everything, but you got to decide where you start.
And so for that team initially, we started really like focus on like metrics and just getting more metrics into like what is going on?
What is the impact?
What is the cost as well?
And on the other side, we knew already that AI was pretty expensive.
And so we wanted to start to be able to benchmark different models.
What is the best solution?
And so we created a team or that team started to work on evals and building kind of an agent decoding tool, eval platform, to help us make decisions about like...
which model, which tool, which is on board, I did it up.
And this is, so this is like an early embrace of evals, which I appreciate, and we're sort of big fans here at TESO.
And how did you apply that?
So like, what are examples of things that you'd eval?
We'll dive in more, so I'm still in like journey mode, but these sort of earlier days, it was sometimes hard, like you might draw a conclusion from an eval, but it wasn't, it's not always easy to think about how would you apply that learning.
So what are examples of, like, should we do X that you've decided?
Yeah.
So I'd say, like, the first application of evals for us very concretely was for code review.
And we had a bit of back and forth on, like, you know, what's the value of AI in code review?
And I think, you know, if you work with a tool that is very nitpicky, you know, it's not really fun to work with a nitpicker that is not human.
And so...
And so we saw pretty early on the like, really for us, the distinction here would be like, we wanted to put AI in the loop for code review to really like prevent an incident or like kind of major bug and like really like focus on that, you know, as like kind of the last guardrail before you go to production.
And so the first evals we did was like replaying PRs, we knew had caused incident and then like catching, you know, are they...
gonna figure out that this is gonna fail actually um and so you know we do track a lot of data around incidents here at datadog in a data set and so we had all that data already kind of there for the picking and so we were able from there to like devise and like kind of basically built a platform that would replay prs and then run the whole agent review and then with an agent judge being like well does that does that review, comment, point out what is going to cause an incident later on.
Yeah, super smart because it's like a clear use case because in evals oftentimes the difficulty is defining what correct is, but if you have those historical incidents.
So I want to dig into evals, but let's just sort of complete maybe the timeline.
So you went ahead, you sort of hired or like moved this internal person and you started building a team.
And I know, you know, maybe fast forwarding to today as that team evolved, you now have two teams, right?
Like now that you're sort of a bit more mature, maybe tell us a little bit about current setup and then we'll move on to evals.
Yeah, yeah.
So what it says, like we had a lot of ambition, you know, around like changing the way we work with AI.
But towards the beginning of the year, what we saw is like that first team that we built, they had a ton on their plate with evals, with metrics, with like cost management and all of that.
And so it was really hard to then like, you know carve out a piece of you know time for us to actually work on like rethinking how we work and at the same time at the beginning of the year i think everybody saw with opus uh four five uh that came out just before the holidays the like adoption from there really like took us like really really fast you know probably everybody says like the the amount of pr pretty much doubled overnight you know yeah um and it started to really show you know kind of bottlenecks in our flow in the SDLC that are way more serious than they were before.
And then it was hard to carve out time to actually focus on that.
Mainly for us right now, it's really code review is the big bottleneck that we're trying to tackle.
And at that point, we're like, well, we need more people to work and really focus on the experience itself of the SDLC with AI.
and not just more kind of the governance and like signals side of the world.
And so we like launched like a second team.
I forgot exactly when, but towards the beginning of the year.
Yeah, very cool.
I love that.
I like the names you sort of mentioned to them as well, which is Signal and Flow.
It's like very, yeah, that's what it says on the TN type names.
And just, I guess, for it's useful sometimes as people calibrate for them to sort of think about rough sizes of those teams.
Like it's kind of, you have.
circa 4,000 engineers in total was sort of the rough sizes of these two teams.
Yeah, yeah.
So we're aiming to have six people, around six people in like a boat team and we're still recruiting for that second team.
Got it.
Yeah, that's still a good ratio, you know, in terms of multiplications and leverage.
Yeah, yeah.
So about, you know, 12 people for 4,000 engineers.
Yeah.
And the, I guess as the organization, flows, this may be like one last question on the team thing before we continue.
As you change the way that you work and the adaptation, did a need for new signals come out?
Because I think the previous team, like the signal team, is very much around current way of working, let's sort of measure output so we see if productivity improves or costs change, etc.
But now that you change it, did you find the same metrics work over time or did you need to?
kind of invent new metrics, right, or sort of start measuring new things?
Yeah.
So I don't think we've had to change dramatically how we measure, you know, productivity.
But I'd say, like, we're using a lot, like, Dura metrics, and these tend to be more, you know, they're not, like, leading metrics, they're a lagging metric.
And so looking more into, like, the leading metrics and, like, what's in there, I think usually we need a bit more creativity and those are usually kind of less like true signals, like they're a piece of the pie, it's hard to isolate exactly what moved those metrics.
So what I'd say maybe what has been more visible is more like how we think of cost and the amount of token we're using.
And then just being more professional and handling cost of AI.
Now everybody at the company basically using AI, using more of it, there's a bit, you know, it does increase for the budget that goes towards AI.
And so like, we just need to be more cognizant of this and just make sure that it's just used reasonable.
And so, and so definitely more on the cost aspect of it that we've done.
And then the other thing is like with evals and now that we're, that we're getting a bit more data, then like the question comes on like, well, how like do you turn that into action or like to know?
what to do with the evals.
So it's like, probably like those two aspects had changed more in terms of productivity metrics themselves.
We're still mainly looking at like Dura metrics overall to get signals.
Yeah, yeah, interesting.
Yeah, makes sense.
And I guess, did you ever get into token maxing?
Or has that been, because you mentioned cost, like you're sort of aware of how much it cost before.
That sounds a little bit counter, like I want this successful, but I want it responsibly successful.
No, no, not like token maxing.
We never did.
quotas for example uh and so and so we're trying to like really not go that route because you know if you need a lot of token because of what you're working on you know we kind of do want people to have the optionality you know and be intentional about for their for their own use and their own needs um and so you know it's like maybe a bit like breathing hot and cold at the same time uh but we're really more looking into like systematic way that we can reduce token or like compress output, you know, things like that.
Cool, cool.
Cool.
Okay, thanks for sort of the run down there, sort of on the team, we've done it.
We've started sort of teasing out a little bit, this sort of investment in evals, you know, and creating it.
And I think I thought that was really interesting, because this is very much like a world of vibes, right?
I'd say even until today, you know, evals, like if you're building an AI product, everybody will say evals.
But if you are talking about sort of development with agents or if you're talking about even sort of like agentic workflows as a whole, I guess kind of showing my bias a little bit is like I feel like there's a lack over there.
Like the vibes are quite dominant on it.
And so I like that you kind of...
created those.
Can you tell us a little bit about what works?
You've invested in evals.
We don't really need to go through the whole journey of, but where you are today, what is it that you do today with evals?
Maybe tell us a little bit about the tech stack, but also what is the value that you derive out of them, and maybe where are they not in yet?
First, it's a fully custom solution written in Go with Synbox.
So I don't know, you know, when we started, if there was a lot of open source option, we pretty like looked at a few, you know, but nothing was like really mature and so like we just went our own way since like it was not like overly complex for a pure use case of like AI agent.
And then Datadog has a pretty good metric platform.
So like it was easy to just feed the metrics in there and then consume it from Datadog directly.
So in terms of what we do today with DEVALs, there's two things that we're looking at really specifically.
On one side, it's just tracking performance of new models, especially for the open waste model, and trying to get a sense of when could that be a viable alternative, or in which fashion could they be?
And on the other side, it's like, how does the...
uh what you know uh things like steering documents um are they are they working uh when we change it does it change the performance of the evals um and so like really more on on the performance side where like teams that are not us are gonna own steering documents or maybe should own the steering documents because we can talk of ownership there sometimes a bit hard within repositories And just tracking a bit the performance on is there an issue?
And should we change it?
And in which form should we change it?
And did you...
Maybe there's two different conversations here.
Sorry, I started both together there.
Maybe one is, let's talk a little bit about the stack, and then we go to the substance.
So on the stack, you built your own because you have your Datadog platform, kind of clearly very good at collecting data.
and aggregating it.
And I guess at the time, nothing else was mature.
Do you think people should build their own eval system?
Or would you say, like, if you were to do this again today, you'd probably pick Harbor or one of the existing frameworks for it?
Or do you think there's sort of aspects of how you do the eval that you feel like they need to be custom, they need to be yours?
Yeah, so don't feel like strongly about custom versus hosting platform.
I think...
what's important is to be able to support multiple harnesses, multiple models from a different source.
I think it's pretty important.
But as long as you have that, and I think most of the platform do offer it, I think going with an off the shelf solution is pretty good.
But I'd say also coding being less of an issue today, for us, it's been fairly easy to put it together.
It wasn't that hard.
Yeah.
To build the system.
So I guess, so I think that that sort of makes sense.
Like technically just to be able to run an agent and sort of have a card or like some sort of scorecard and run that and aggregate the data, you know, those are kind of within core competencies.
And I guess you kind of always weigh that a little bit against, okay, but why reinvent it?
You know, like if you, if it's not that custom, so it's a decision, I guess all developers are sort of fairly familiar with.
The second question that I have is maybe a bit more in the substances.
At the beginning, you talked about use cases that made a lot of sense.
that are central and have clarity about what the truth is.
So I guess my question is, as you approach use cases that are sort of more elaborate and talk about using models and all that, like what is correct is not always as clear cut.
So I guess how do you handle that?
And maybe like who even writes that down, you know, today, right?
Like who creates the eval scenarios?
Yeah, so in terms of evals, And like writing those evals, there's a few different things that we're doing and like new stuff that we're trying out now.
The status quo so far has been we've wrote the first ones with like very common use case to just like have a baseline that we can go through.
And then we went out to pretty much like all of the platform teams and like ask them to like add evals, you know, to like cover their like own platform use.
and to really spread out for the ownership of those evals across the company and just get more support.
The downside of that, though, is that, yeah, it is limited in a way of like, if as a dev you're using like a harness or like an agentic tool, like you're probably not only asking them to do very specific things.
There's a lot of the work that's more on the creative side, finding ideas, what should you do?
What's the right approach to your problem?
This we don't try to cover with the evals we have because it's just so broad that it's just really hard to know what exactly should you try.
And if it's always the same question, then you only have this small slice.
But the creativity aspect, I think, is somewhere or something that is really hard to gauge with an agent.
So I don't have a great answer to that.
But I know it's a hole that we're having and we're fine to live.
with this gap that we're not sure how to fill yet.
Maybe someday I can come back and maybe we figure it out.
But for now, we're not too much trying to tackle that.
So you just sort of echo back over here.
You're saying at the moment where you're sort of seeing the value in a bit of a...
concrete fashion is in finding cases that have clarity about truth.
They're repeatable processes.
You're doing the thing again and again.
And generally, the scenarios and all that are written fairly centrally, like not just by you, but by, you mentioned, platform teams.
So it's not really going to the end developers.
They're not really writing evals.
These are extractions out of real-world scenarios of evals that change at a slightly lower pace.
broadly speaking.
And you think eventually, like if you were kind of a betting man, right, and you sort of say in a year's time, do you think all developers write evals or do you think it stays this way?
This is like a crystal ball question, so no, we're not going to do it.
No, but I have a good answer for you here.
I don't think every DAB is going to be writing evals.
Though what we're doing right now, so, you know, TBD, if it's going to work, but I'm pretty confident that it's going to be the right solution for us going forward.
We are tracking...
for the trajectories of agents and what they're doing.
And so we're tracking a lot of data today for the use of agent within Datadog.
And so what we want to do is like go from that raw data of like, you know, well, this prompt came in, the agent did that, et cetera.
And like extract from there common scenarios that tend to recur or like common type of interaction that a user has with a given platform.
and then using an agent to turn that into actual evolve.
And so basically taking like real data of like how people use the agents and make sure that we're covering, you know, a fair amount of overlap with actual use.
And we're pretty confident that we can automate most of that.
Yeah, yeah.
No, I love that.
That is actually very aligned with sort of our experience at TESOL, which is the more you anchor it in concrete cases.
the more real it is.
And then, you know, like, what the developer is not like doing, they don't like writing docs, they don't like writing tests.
And so now you're sort of saying, hey, write docs and tests, you know, that doesn't really work.
And so probably the best ease of use aspect to it is instead of asking devs to write those tests or write those eval scenarios, you observe them, you sort of extract them out of the loop, and you extract these scenarios.
I guess there's still, there is an unavoidable...
future sort of evolution to that that you will still need like these things are dynamic and so you still captured that scenario that sort of hypothetical aggregate PR at least that's the case in our world right you've observed a bunch of PRs you've created an eval scenario you're running it you're still capturing that as some files that sit in some repo and over time it will rot, right?
Like over time, it will no longer represent your system.
And so there has to be some maintenance.
So at some point, someone needs to own that, right?
And that someone is hopefully someone who can kind of represent developing in this repo, right?
Or something like that.
Yeah.
Yeah, I agree.
But I think it just, you know, it's hard to catch rot and it's hard to come back to, to like case you created in the past because like amount usually, that you end up with is like fairly large and so we're really focused on uh on automation here so like really having an agents to to do for the sweeping at the release bring up the signal but eventually maybe even make the decision about like which case are still relevant today yeah yeah i think uh I buy into that.
I think in general, whenever you introduce a human, you create a bottleneck.
We'll talk about that in a sec.
But whenever you introduce a human, so you have to plan ahead for this.
And in that sense, the users, both the authors and the users of the evals, are very likely to be agents.
So you still need to create them, but it's not necessarily the person that creates it.
It might be someone else.
And I guess the...
before we kind of move on a little bit from evals on it, how often do you run the evals?
Like once you agree that something is correct, is this something you deem after the fact?
Do you run them on every context change?
Do you run them?
How does the execution of the evals integrate into this SDLC?
Yeah, yeah.
So today we're doing Nike evals.
So we're running basically once a day.
And then for PR that might specifically target evals, like they're changing steering documents in a heavy way, really trying to change the results.
For those people who can do an ad hoc run, and they're probably going to want to do a few since evals are not always super stable to get statistical significant results.
Got it.
Very cool.
So this is, you run nightly.
You look at relevant changes, some steering docs that have sort of modified.
If those occurred, you run the eval.
Even if there's an apparent regression, you point that out to the person.
So you don't get in the way, and you run them.
I guess it's sort of consistent with the previous theme, which is you run them on the things that you feel like you truly know as opposed to trying to speculatively anticipate ahead.
Correct.
And that probably results in something that like cost wise is not that big a deal.
How expensive are these evals getting?
It's a good question.
I'm not sure, you know, but evals are expensive.
They're cheap enough that they haven't doubled.
Compared to the coding agent use.
Yeah, it is nothing compared to, but it's definitely something that we're cognizant of.
And like the decision to run it likely is obviously part of that.
We're like, well, if we run per PR, obviously it's.
that is going to be too expensive.
Yeah, it's interesting.
We've been experimenting with this, of course, in Tesla, like Tesla itself, but also with customers.
And one of the observations is that you can actually use cheap models, like cheap open models, to assess regressions.
So we found that DeepSeek actually correlates pretty well in terms of the skill uplift for a variety of tasks.
And so what you can do is it doesn't need to be as good.
at overall the tasks, but if you measure it with skill, now that one is like a lot, it's fast and it's cheap.
And so maybe you can run that sort of DeepSeq V4 flash as part of your CI test for some sort of selection.
And so you find it, because it has all the sort of the natural problem of if someone made a change yesterday and you're coming back to them tomorrow, it's annoying, it's rework, it's better if it's in line, have to weigh that.
So I don't know precisely where it lands, but I feel like there is an evolution of evals.
expanding as tests, right?
Just like we have some tests that we run in the CI and some tests that are too expensive or slow to run in the CI, I think we'll have the same for you both.
I think that is exactly correct.
And it's also the way we're thinking about it, right?
Like what you can run after the fact, unlike, you know, what's the cost of a revert and like a rollback versus the cost of like, that like you pay now waiting before you can merge something.
And so it's not always a clear trade-off.
But we need to try to be cognizant of the cost of making a mistake versus the cost of waiting for things to run or paying for the info to run.
Yeah, exactly.
Very cool.
Okay, cool.
So I love the passion of evals and I can kind of geek out about evals.
I've built a passion for them.
But let's talk a little bit about context.
So in the process of that, you referred here to sort of steering context and all that.
What is the state of context typically within Datadog?
Is there a central context where people share it?
Is it mostly in repo?
How would you describe skill adoption right now?
Give us a bit of a picture of where is context being used.
Focus on dev within the organization.
Yeah.
So I'd say there's two big sources.
One is files that are just committed to the repositories themselves.
I'd say skills is now maybe what there is.
There's nothing.
a lot of other stuff that is not skills skills is pretty well adopted.
Yeah.
And then we had the cloud marketplace with a lot of skills there as well.
That is centralized.
And I can talk of like mistakes and like all of those as well.
If you're interested.
Let's talk.
I think you had that sort of insight.
So yeah.
Yeah.
Let's start.
Let's start positive.
Where did we're used to see it working?
Yeah.
Yeah.
So, I mean, what.
What worked really well is everybody contributed early on, and especially when we started without a central team, it was like a PC just to get people onboarded.
I think it was good to not put too much structure early on so we can try and see what worked and have learnings out of it.
And so I think for the barrier and the friction early on was really low, and I think that was probably the right move to do.
because we want people to be contributing and we want the best idea to win and those ideas to be easy to share.
Yeah, I love that.
The kind of core principle about transformation is at the beginning you're trying to drive adoption on efficiency.
Just sort of get people in the boat.
Absolutely.
But then what we saw through time is the more user we've had, the harder it became to go and change them because the impact you're having then, you know, instead of being on 200 people doing a POC which you're using for the same tool, you're going to impact 4,000 people working on the repository.
And so it's just much more stressful to break things for so many people.
And then also the size became kind of hard to manage.
The cloud marketplace has hundreds of plugins.
once you install the marketplace, like, do you know which plugins are good?
Which one should you use?
Which one are going to perform better?
And so what we've steered people to do is more team based marketplace for them to share for their own flows within the team, things that are specific to their team or marketplaces extension for specific platforms.
So if you work with platform X, well, we'll So we will suggest in the documentation that you install that marketplace and those plugins for you.
Or we'll refer to it from context document where the agent can suggest to install a marketplace and stuff like that.
Yeah, very interesting.
So like groups of contexts, so people know where they're approaching that and they control that.
I think this kind of slightly hinders...
maybe like reusability, you might sort of have different teams.
You probably do have different teams that do that, but you slightly control the chaos in this fashion.
Yeah, yeah, yeah, yeah, correct.
Correct.
And then what I'd say we found out at the beginning of the year, Pi became like pretty popular recently.
And so with so many people like starting to hint at like the fact that like Pi somehow was just performing better than the alternative we had.
And Like, clarify for those who don't know that Pi is the very highly customizable custom harness, but it's open source, but then it can customize on top of which OpenClaw is built.
Exactly.
And so one thing is like, by default, it's not going to be the lot of the context.
And so it's kind of the hunch we've had is like, maybe actually, the context that was written like a year ago, you know, when it was like Purser, Sonnet 3.5.
Is it still relevant today in the world of Opus 4.8, of GPT-5.5, right?
Like, maybe not.
And so we started to look into it.
And then we're like, well, what if we delete the whole context from the repo?
Like, what would happen?
And what happened for us is actually, like, the evals started to perform, like, much better once we deleted the whole context, which was counterintuitive.
But it's just like, yeah, that file was pretty big.
and was listing stuff that was not like necessarily useful or, you know, that should not go into like a steering document.
Yeah.
And so that was really interesting to see.
I love that sort of insight.
It's kind of a real world manifestation of how context rots, right?
And we're sort of so familiar that software rots that it goes from kind of being useful to being useless to being harmful on it.
And for context, we don't have that muscle memory yet on it.
So now you're seeing it with the skills.
I guess what was the result of, so the insight is great.
It sounds like you still ran those evals centrally because it was like you or your team is kind of thinking you'll do that.
You ran the evals, you found out it's actually outperforming having context.
I guess what could you do about this?
Could you sort of go off and say, cool, let's delete the context?
How did you handle it?
Yeah, yeah.
So I think if we didn't have data to back it, it would be pretty controversial uh and just like to to be clear this was owned uh by the front end team that run uh for the front end moto repo they wrote eval they did the folks for the for the testing there uh and like kind of they wrote new guides like what should you do instead of context and so like you know that was fully something that like was owned but by the sort of team within my org you know but but but like not like by the ai team uh and so and so yeah uh With the data backing this decision, a lot of people were surprised by this decision, but the data was really clear that it was performing much better.
And so people agreed with it very easily once they saw the data.
So I think two sort of interesting tidbits over here.
One is you removed the context, but it's because you have started forming a harness.
So you're saying this was not just because the models got better.
but also because you started working with more custom harnesses?
Oh, this is something we're doing now.
So this is still primarily to that decision.
Yeah.
Okay.
So this, I guess it's sort of still primarily either the model improvements or just those context files were created at one point in time and they no longer represent reality for the agent.
So anywhere between the new models don't need them and they're just wrong at this point.
Correct.
Correct.
I don't think they were wrong per se, but I think it's early on.
There was a lot of basics that he would try to retrain the agent on.
Yeah, so in the case of the front-end repo, there was a whole paragraph on how to use YARN.
And I'm pretty sure now it's part of the training set of the models out there.
And so we don't need to retrain it.
Yeah, yeah, that makes sense.
And context is a scarce resource.
And so over time, if you accumulate enough of that, you take away.
I love that I saw one talking where someone talked about how the model gets into dumb mode.
Yes.
Okay.
Cool.
So in terms of governance, you know, around context, if you want to go a bit in there, is what we saw is like, actually, like, I think we have an inclined to just go in and add something for like an agent CMD.
But very often, it just has easy as like, is the error message from that tool clear?
And do you provide the next step of like, what should you do instead?
And so we've been also moving or like pushing people that like, you know, when they run into it, like kind of an issue where the agent doesn't do what you want it to do.
Well, where can we provide that context when it's necessary instead of like trying to train the agent ahead of time into how to do everything, especially within a monorepo where like there's so much going on.
you just cannot provide all of the context up front.
And that seems to be detrimental.
Yeah, yeah, no, very good.
So I think a whole bunch of lessons from here.
One is the importance of evals, you know, so you can make data-driven decisions.
You know, two is maybe a reminder around how context does rot over time.
So you have to have some sort of plan around how would you sort of reassess it, which comes back to evals.
But even if you're just vibing it, like still do something, like be proactive around updating it.
And then, you know, maybe just a little bit more as you write new context, solve problems, don't sort of anticipate them as you go, which I guess comes back to the observability and the loops, right?
Because today, probably the best practices don't actually solve problems, like just sort of solve a problem in the spot and create a loop to identify that context.
And you, like when we spoke, you talked about how...
you think one of the theories of how the context got accumulated in the repo was because of this sort of loss aversion.
Like, you know, once someone wrote it, it was hard.
And maybe this comes back to evals is, you know, it just sort of feels fearful to delete it.
Like maybe it is useful somewhere on it.
Is that, do you want to?
Yeah, you know, absolutely.
You know, I think like as people propose to remove stuff from like that root agent MDFAR, like, Like even in like PR review comments, people are like, well, are you sure that like you should be doing this?
Because like maybe it provides value to somebody somewhere.
And so I think, yes, it's when you remove something and you don't replace it with something new, it is just hard to do, especially when you're having an impact on like hundreds or thousands of engineers at once.
It's like it's hard to have certainty.
And so like if you don't have data to back those decisions, it is really hard.
to take the step to go ahead and do it or feel the permission or feel that you have a good reason to do it.
Yeah, yeah, yeah.
I'm 100% sort of in agreement and I love to see it manifest.
So I guess in those marketplaces that you have in the small ones, coming back a little bit to what do you create evals for, I mean, this sort of logic will imply that whether now or soon, are you starting to feel that within each of these teams' marketplaces, they need to have evals for how?
how well these skills like for the things that are shared like it goes to reason to say like they also need devals because otherwise they'll have the same problem yeah uh yeah yeah for sure uh they will i think like maybe one of the challenges also just like the distribution um of those skills does everybody have them installed when they need them and so there's a bit of a challenge there of like making sure that people have the right tool for the task that like they want to do And I'm not sure the way we're working today with local agent only and everybody set up for their own thing.
There might be limits to that eventually.
And so I'm not exactly sure what is going to be the right solution going forward.
But I think meta harnesses, for example, are pretty interesting to that of like, you know, what if we have kind of a classifier and we know, well, this is touching platform A.
So let's...
to run this with an agent that has those skills pre-installed.
And if this was centralized, for example, I think that might be able to provide a much better kind of experience where we have a better baseline that don't rely on tribal knowledge or somebody having read documentation.
Yeah, I agree.
I think generally people underestimate the importance of environment definitions and evaluations.
And so I think that's at least how it translates in our world.
We created these project evals versus skill evals.
They're actually all just the same evals, just names.
And the idea was the skill eval is a thing you test in isolation.
And so it's a good way to test if it regresses or not.
that needs to tackle that first problem, which you correctly pointed out is a bit hard, which is define what correct is, because it runs together.
And then you get into the problem of like, okay, that's nice and well, but I actually have 30 skills installed when I'm running this, not one.
And so does it work in that environment?
And that basically gets infinite, right?
Like which skills do you have at what time?
And so I think environment management, like the thing that we've sort of seen extend in flexibility, but then it's also in complexity, you have to think about how to rein that in.
It's just like all these definitions of how do you define the environment within something.
runs.
And I guess they both boil down to trying to eliminate works on my machine scenario, right?
Like the one is define what correct is and define the machine.
Like define works and define machine and you can step out of that.
So it's interesting.
And it sounds like that's still sort of a work in progress.
Like at the moment in practice, these little marketplaces, I don't know if little, but like team marketplaces don't tend to have emails.
Yeah, correct.
You should do that.
Yeah.
So these are fascinating conversations.
We started to run out of time here.
So maybe sort of chat a little bit about kind of open models on it.
Maybe also open models and bottlenecks.
So you started by sort of talking about how the evals were useful to access models.
And you've alluded in our previous conversations how today the bottleneck is very much in the review phase.
of it and where we work over there.
So I guess maybe I'll kind of ask you a little bit of a joint question, Abed, which is, you know, where are you and what are your sort of views around the review bottleneck, you know, what to solve it, and have you tried open models on those?
Yeah, so yeah.
Based on our evals, I think open-weight model to be like fully viable, assuming the bar does not move as progress happened in the space would be that they would need to perform about 50% better.
than they are today.
Okay, interesting.
This is a statement even for like JLM 5.2 and like the latest or is it a notch below?
No, a notch below.
We didn't yet run the evals in those, but we're very excited about this.
Fresh capacity like these dancers every time.
Yeah, yeah.
Yeah.
And the caveat being like also for the creativity of those models with like tasks that are unknown or, you know, larger tasks.
instead of the execution part of it.
And so this, we don't have signals from our evals, but we're looking at public evals to get some signal there and then getting user feedback.
Yeah.
So at the moment when you run, say, code reviews or these highly repeatable processes today, you're still with the frontier models running those.
Yeah.
For code review today, we are.
We're having the background agents.
that for example for code review the first one we release is like if your CI fail uh you know and it's like a linting error a formatting error stuff like that and like that's pretty easy to offset to like an agent to just go fix it for you so you don't need to like come back and like have to deal with it yourself and so we're trying to like remove friction from from like different places and unlike a lot of those tasks can be done because uh well by like like an openweight model that is like much cheaper because it doesn't need creativity.
It's like a very clear path.
Like you need to run a command or you need to do some kind of something specific.
And so our thinking now is like we probably can't encode a lot of what you would do out of Claw code to background agents that are just going to take action on your behalf to like nudge things along.
And then here is like finding the right balance between like what can we trust, you know, and like.
not changing core decisions of the PR, really more into nitpicks and small fixes without changing the meaning of the code change that is being proposed.
Yeah.
And this is consistent.
It sounds like your view is there's a whole bunch of places in which it's inarguable that the agent helped.
Let's apply the agents in all of those places.
Don't get into the argument of whether it's right or not.
Get it into all those places in which it can do it.
And clearly, there's always like a little bit of a halo, like, you know, you might veer a little bit into judgment, but mostly lean into the areas in which you can inarguably say, this is correct.
So can the agent do it?
Can it not?
Does that sound right?
Yeah, that sounds about right.
Yeah, yeah, yeah.
And I think it's like super practical and kind of a high impact.
And it sounds like in review world, and I guess kind of in the repo world as a whole, you're just sort of looking for these, surrounding annoyances to sort of resolve those.
And there's so much kind of meat on that bone that you're plenty busy sort of doing that.
You don't need to go into features quite yet.
Yeah, yeah, yeah, yeah.
Correct.
Hey, everyone.
Hope you're enjoying the episode so far.
Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development.
whether that's talking about the latest tools, the most efficient workflows, or defining best practices.
But for whatever reason, many of you have yet to subscribe to the channel.
If you're enjoying the podcast and want us to continue to bring you the very best content, please do us a favor and hit that subscribe button.
It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you.
All right, back to the episode.
I'd love to talk a little bit about people and around hiring.
So I know you've had a lot of learnings about how you hire and what is it that you advocate or expect people now when you think about promotions or their career ladder?
So can you share a little bit about the learnings there and where you are today and maybe some key learnings?
Yeah.
So personally, I always hated lead code type of interviews.
where I think like they were low signal, a way to like filter out some like candidates.
But I think like very often, you know, could just lead to like the wrong conclusion.
So I've never thought that it were a nice way to hire people.
I've never liked doing them personally.
And I've personally always felt like it was kind of low signal basically.
So I think AI opened the door to make a case to actually like review how we do.
interviewing especially coding interviews um and so um and so what we've done is is like we've built a few squads to own different questions um and with ai in the interview it allow us to do much more than we could before and so it allows us to basically work in a much larger co-base and asking like actual tasks like somebody might be doing in their day today And like assessing more signal than like, yes, they could write the code in like 40 minutes and like pass all of the tests, but more like how's their judgment?
How's their trade-off?
Did they understand the code?
And so like with AI, basically the new flow for most of the interview is going to be like, here's a code base.
It's fairly large.
Use AI to understand what it's doing, what is going on, what is this code base about?
And then afterward, like, let's like talk of an engineering problem.
So like, maybe you want to extend this code base.
Maybe there's like a diff, like basically kind of a code review that like somebody wants to merge.
Should they, what should we bring back in terms of like feedback that I could review?
Yeah.
And so more like kind of real world scenarios.
And I think this can only happen with larger code base.
And without AI, it would just not be possible to go through and understand the whole context.
It's like it would be more left to luck almost.
And so I think it really allowed us to step away from lead coding into more actual real-world scenarios in the interviews.
Yeah.
No, I love that it, first of all, it's closer to the reality.
And there's actually a bit of a hidden question about your sort of agent foo.
You know, like how magical are you at kind of navigating the agent?
you're sort of testing two things at once, you know, your depth in the agent usage and your ability to understand those.
And I guess it does imply, though, that this is like an in-person type element.
You know, there's no, this isn't a kind of take-home code review or like coding exercise.
It has to be inside.
And is that, you just sort of accept, I guess, that's a slight inefficiency in the sense that you need to talk to more people that you just take this with.
Correct.
Correct.
And how, I guess on the other side of it, like when you think about people moving up the ladder and you think about people already in the job and doing it, what has changed in terms of what do you assess them for?
Yeah.
Yeah.
Yeah.
So we reviewed our career guideline this year too.
And so we got together with a group to like brainstorm about like what change in terms of expectation, like what is more kind of important today?
And so a few things to call out at the high level of things like we changed.
One is for more junior roles.
I think we saw for everybody that we had a higher expectation about the complexity of the work that they're going to own and push forward.
And so we updated those roles to actually reflect the new baseline expectation of just like, yeah, you can take a whole project.
you pretty want some help from like somebody you probably kind of understand asking the right questions um learning about you know the uh sort of core competency of like being an engineer uh but compared to before which was usually somebody would be in a support role on the larger project we're finding that like most of the time now they are they are owning like kind of a whole work stream by themselves because they can now uh and then they just need to make sure that they're asking the right question and so like Here we change a bit that definition for more junior roles.
In terms of more senior roles, I think what we saw is like before, well, two things.
One is in terms of like POCA and experimenting.
And the second is like on like judgment.
You know, like, should we do this?
When should we stop doing something?
Is this solution the right one to the problem the customer is like having?
And so like there's much more of like a judgment call being done.
Yeah, and then in terms of...
Product and architecture tastes, right?
Yeah, yeah, correct, correct.
And on the POC side of things is that I think before we would spend a lot of time on documents and RFCs to talk about, is this the right solution for that problems and all this?
And I think there's definitely value to those questions.
But I think now we can know...
more certainty around like is this gonna work or not if you do a poc first and so you can validate assertion in practice by doing poc and because it's so easy to like poc something with ai is like you can do a few poc and like compare those options before you decide into like which one you want to actually invest in long term and like really make production ready and so it's like we have much more of like of like kind of a try things first and then bring the learnings from that to make the proper decision.
And so that's a bit different from before where you would need to think of everything because it was costly to create a POC.
But today, yeah, try something, run it and see what happened and then bring those learning back to then inform the right decision you want to do.
Yeah, I love that both as a behavior and as an expectation from people.
And I often say that now that the cost of building is lower, the relative cost of alignment is higher because it takes just as much time to get people in a room and have a conversation.
And so investing and building more to reduce the cost of alignment because by the time you align, there's something concrete and kind of more proven is clearly the right trait for it.
So I love this.
Yeah.
And I think it helps a lot in discussion to just have something more concrete to talk about.
You might have concern, oh, this is not going to work.
But if you had it in a PUC and it does work, It might prove the point that, okay, no, this is possible, actually.
So, like, give you a bit more certainty in your direction.
So, I think, like, it changed with you the whole game into, like, discussion of, like, what's the right design for a solution.
Yeah.
So this was excellent.
It was like just jam packed with useful advice.
I was going to ask you, what's your sort of single biggest lesson or takeaway?
But frankly, you sort of shared a whole pile of useful lessons here and learnings for it.
So really, really appreciate you sharing it.
I'm sure the listeners are.
Instead, maybe I'll ask you like one different question, which is what are you excited on for what's next?
You know, what are you looking for forward the most?
Yeah, yeah.
I'm really excited to...
basically like being able to make better business decision and like that's kind of where we ended up on but like that's really my belief around ai like productivity is nice and all this you know but it's not like really kind of a game changer like but yeah being able to know and make better decision i think is really the the key thing where we are like gaming with ai and this is like really what i hope we will see happening in the business world is like hopefully we ship better product because we were able to try multiple things and see what works and what doesn't instead of like maybe spending kind of a whole quarter on a project that is like off ready, then you ship it anyway because of, you know, you've like invested all of that time.
You need to show something for it.
And I think that like calculus here of like what can go out and like when you can be like, this doesn't work, it's like start from scratch and do something different with the learnings we've had.
I think it's much more possible today than it's been in the past.
And so I'm really excited about that.
Yeah, that is super exciting.
We did a roundtable at AIA Native DevCon recently and someone said there that it's not about increasing productivity, it's about increasing ambition.
And I loved that sort of phrasing.
It sounds like that's very much what you're talking about here.
I love it, yeah.
Cool.
So this has been excellent.
Thanks a lot for coming on to the show and sharing all these learnings.
And for our listeners, thanks for tuning in and I hope you join us for the next one.
The AI Native Dev is brought to you by TESL, the package manager for skills and context.
Your hosts are Guy Pajani and me, Simon Maple.
Our producer is Tom Dowler.
The AI Native Dev is not just a podcast, it's a community.
And we host monthly meetups at the TESL offices in central London.
Visit tesl.io forward slash community to learn more.
And I hope to see you there.
