# Intercom's Agent-First Engineering Transformation

**Podcast:** Engineering Enablement by DX
**Published:** 2026-06-15

## Transcript

Welcome back to the Engineering Enablement Podcast.
I'm your host, Justin Riach.
This episode was recorded live at DX Annual, where we were joined by Brian Scanlon, a Senior Principal Systems Engineer at Intercom, who helps lead the company's push towards agent-first software development across a large legacy code base and global engineering organization.
In this session, Brian shares how Intercom doubled engineering throughput in nine months and tripled pull request throughput over 16 months by standardizing on cloud code, building hundreds of internal AI skills and restructuring engineering workflows around agents.
He also dives into the operational realities behind AI adoption at scale, from auto-approved pull requests and AI-generated code reviews to organizational change management, tooling strategy, and the real cost of running AI native engineering teams.
Thanks for listening.
So great to be here.
I've had a great day.
I'm like, this is the coolest room I've ever given a talk in.
Such an amazing view.
Yeah, so I'm...
from Intercom.
I've traveled over from Dublin to Ireland to talk at this conference, but also I kind of feel like I've traveled from the future, like looking at many of the talks and from telling to other folks over here on my visit, I think Intercom is a little bit in advance compared to where most people are at.
Although everyone is pretty much working on the same stuff as well, which is really interesting.
So I'm going to be talking through what we've done, how we've achieved doubling throughput of our engineering team.
And it's going to be pretty open and honest with like some of the information, things that didn't work.
But I will be talking from like go through some thought leadership, the engineering leaders sort of content.
But I'm going to go to some detail into like the actual cloud code skills and things like that that we use.
And so Intercom is a 15 year old Irish American B2B SaaS startup.
And we've pivoted.
kind of like everyone else, but we've done it pretty successfully and aggressively towards being an AI company.
So this graph here, this is like the growth rate of SaaS businesses, like kind of in our cohort.
And you can see like SaaS not doing great generally, but Udicom is doing great.
We've done like turned around growth rates and like we're certainly booking the trend that SaaS businesses have been suffering from recently in terms of growth and valuation and things like that.
Additionally, we've been a bit of a poster child for companies redefining themselves in the age of AI.
New York Times recently wrote an article about SaaS companies reinventing themselves that prominently featured Intercom.
And being an AI company means a lot more than just slapping together some wrappers around your existing product.
We built an AI agent for customer support, completely displacing, effectively, our previous products, which is a help desk.
We've got 8,000 customers, 100 million in revenue.
Growth rate is extremely positive, you can see in the previous growth metrics.
And companies like Anthropic, Snowflake, Linear, LaunchDarkly, Glean, they all use FIN, our customer support agent, to support their customers.
And we recently announced that we have our own model serving 100% of FIN.
These are like displacing use of the frontier models.
So we've done benchmarks, comparisons, and our own trained models are working cheaper, faster, and better than the likes of Sonnet or whatever.
And we're also happy to sell direct access to our models that are optimized for personal support.
But I'm not talking about any of this today.
I'm talking about the work we have done to push.
for AI adoption and improvements in how much we ship and the quality we ship and basically everything.
So at Indicom, 12 years, I'm in our platform group effectively.
We take care of Indicom's uptime, performance, security, costs, management, observability.
We love these monolithic applications.
So we have these giant Ruby on Rails apps and JavaScript apps.
But also, my team, my groups take care of internal developer productivity.
And we are obsessed with shipping.
This is a nice honeycomb sticker of a blog post that we wrote many years ago.
And developer productivity is something we invest a lot in.
And so obviously for the last few years, I've been spending a lot of time working on enabling use of AI in our software development lifecycle and beyond.
So, you know, given that we've been so AI positive and pivoting our company, unsurprisingly, we've been very excited and impatient about AI changing work and adopting AI across.
specifically like building product, building features.
And we kind of went on a normal journey, I think, like adopting the likes of GitHub Copilot and bottoms up kind of people using cursor.
And we kind of got like some results, some improvements, you know, throughput metrics, different things would kind of go up.
But we were ultimately dissatisfied with the results.
We were aware that like the models are only going in one direction and also the harnesses.
I mean, strong conviction that there's...
huge potential or like that we have no doubt that uh ai is going to transform how uh work is going to be done um and so uh but we we wanted to do something about it and not just wait around so we set a simple goal and we set a goal a bit nine months ago that we would double the throughput um as uh as measured by pull requests like merge pull requests um per head of R and D.
So if we grow the number of people in R and D, we should expect that number will go up as well.
And so we could argue about whether it's a good metric.
I guess every metric is a bad metric once it's a measure.
And peer throughput isn't perfect by any means, but it's also completely reasonable.
And I think you should not be ashamed of picking a throughput measure because just adopting AI and being able to get a lot more done, this should just naturally result in greater throughput through the system.
So, and 2x actually as well, it's also kind of not very ambitious.
We think that doubling throughput is like table stakes.
Like really, we're going to be looking at a 10x or 100x throughput increase.
So, you know, when you connect the dots, whatever, it's like this 2x goal.
Maybe we under kind of sold it or we could have gone something a bit more aggressive.
But anyway, we picked the goal.
We also were fortunate enough that there was a big inflection point.
It's been referenced in a few talks here today around the time, like December or so, with the new models and this noticeable shift in model capability and tooling just overlapped with some of our 2x push here, which was pretty nice.
But change is hard.
And it's not just a case of setting a goal or picking the one tool.
There's no one thing that will get us to success.
And so we changed a lot of things.
We took a lot of action to make sure it was very clear to people what we were trying to do.
And so we updated job descriptions and expectations of engineers and designers and product managers.
If you weren't using agents in your work, you are not meeting expectations.
So that's like the stick.
But we have plenty of carrot as well.
Like we reward and promote and through spot bonuses, through like social kind of proof, we praise people who are like acting as like true flywheel creators and helping others kind of be successful.
We give space for people to...
grow and to try things out we did hackathons and enablement days we staffed this stuff full-time as well like you need to be able to support people um so that they can do great work and get over all of the kind of initial uh barriers that people might run into or or just like day-to-day kind of use of the tools they're still they're unstable you know you need to do a lot of work to like get these things working well um and you know leadership is basically talking, or staying on message, saying the same thing over and over and over in every single forum.
And we did that.
And we were very clear about what we were trying to achieve and that people couldn't hide from it or only kind of half-ass their adoption of AI.
We also picked a platform.
So we standardized on a tool.
In our case, it was Cloud Code.
I don't think that was controversial, and it doesn't sound very controversial.
But the main thing is that you pick one tool.
It doesn't matter what you choose.
We happen to pick Antropics cloud code and we think of it as not a tool, but a platform.
And by that, I mean, you know, to get the most out of the tools, you have to get, like, let go of model anxiety and comparing which one's better and all that.
And like, that's sort of interesting.
You want a few people who are doing that in your org.
Most of the benefits are not from the actual models or the harnesses themselves.
They're from all of the context and information and domain-specific knowledge and guides and skills and all of that stuff, all of the glue work, which is specific to your environment.
That is the stuff that unlocks huge levels of performance, accuracy, and productivity with these tools.
pick one.
And we still have a few people using cursor and that, but everybody knows that we're investing in Cloud Code being the platform and building out skills and things like that.
And our vision is that we want to treat Cloud or build Cloud with all of the appropriate skills and everything, such that they can take on any work as a senior engineer.
So this means connecting it up to absolutely everything.
onboarding it, teaching it about everything that's specific about our environment, Rails conventions, React patterns, testing standards, whatever, and train it.
Train it as if you were onboarding a new engineer into the environment.
And so we think that adding all of these skills and all these capabilities, when combined together, will be able to take on any bit of technical work that engineers are doing in Intercom today.
And we want this platform to be self-updating.
We add this into our guidance and skills and stuff to make sure that if new knowledge is discovered or something novel kind of happens when doing a task, update the skill, update the knowledge, and try and get these things automatically.
So yeah, we strongly believe that all technical work is going to become agent-first as quickly as possible.
all of the stuff, the gunge, all of the messy kind of things that we do as part of the shipping product should now be condensed down to this nice, simple cloud is basically doing everything and we're kind of interacting with it.
And we think as well, if the models and harnesses don't improve, they just stay exactly where they're at today, something that's absolutely not happening.
But we have the building blocks that are good enough today to get through.
basically all technical work and make it such that they're Asian first.
And so I wrote down some principles.
These are like a sneak preview of some principles.
We're going to publish them publicly pretty soon.
Because, again, you're trying to convince hundreds of people to change how they work and how to think about how they work and where they should be applying their attention and focus and stuff.
And so these principles, these help people kind of decide and understand what we're trying to achieve.
In this case, all technical work is becoming agent first.
Repeat this a lot.
It means that everything that you can do in your laptop, absolutely every single bit of work, an agent must be able to do those things.
And so having these kind of clear guidance and expectations about how we go about the work unlocks a lot of capabilities.
Or people know that we want them to implement APIs or MCPs or CLIs or whatever to allow the agents to do the work that we do.
We also want people to run less software and that our focus needs to be on the kind of evergreen specific capabilities.
I think the Airbnb talk made a kind of similar point on this.
You know, the space is moving fast.
There's many, many new features and you kind of do need to plug in things like glue work, glue things together or whatever.
But these technologies are moving so fast, you need to like aggressively deprecate.
kind of custom things you do.
And we want people to be focused on evergreen, durable things, which is largely unblocking the agents from getting access to things or knowing what to do then.
So providing skills, guidance, that kind of stuff.
That stuff is way more valuable and durable.
It has lifetime value far beyond, say, vibe coding your own multi-agent orchestrator or whatever kind of fancy things that you can do out there.
It's like we want people to be focused on the core knowledge and tools and access and skills that senior engineers need to do, not on the kind of stuff around that.
We consider ourselves to be leading AI in engineering, and we don't want to get stuck with a custom stack, a load of our own kind of monolith style workflows or capabilities that would cause us to kind of get stuck and left behind.
We also want people to work with agents, not in the kind of telling the agents what to do.
It's more like, yeah, tell them your problems, share your problems with them.
And often a lot of the time at the moment is like, we've written hundreds of skills and a lot of people just tell, like show up in Claude and go like, hey, run this skill to do the thing.
And it's like mostly fine.
I still do it myself.
But really you want to be able to just describe the problem to the agent that the agent figures out from the list of tasks.
list of like skills available and stuff and it figures out what to do and it plans it out and I had like a nice example of this recently where I was paged into a security incident somebody had accidentally published a bunch of snowflake table metadata into a github repository that was public and I just habitually joined a slack channel opened up Claude said hey Claude take a look at the slack channel And I went back and I was looking at the details of the incident and chatting on Slack.
And then like two minutes later, Claude came back to me and said, hey, I figured it out.
And I didn't know that we had a skill that had been defined by one of our engineers, which used our kind of breach criteria and runbooks and policies and all of this to figure out.
it analyzed all of the files that were part of the breach.
It used all of the information that was encoded in the skills from all of our policies and classifications of breaches and that.
And it basically just listed, here's exactly what the story is.
Here's next steps.
Turns out it's a bit of a no-op.
And that removes just about 20, 30 minutes of...
kind of boring work, to be honest, just kind of checking this stuff manually.
It's not exactly rocket science or that interesting, but completely necessary.
But the fun part here was like, I didn't know that that skill existed.
I didn't tell it what to do.
I just like told us that there was a bit of an incident going on and it figured out itself what to do.
And, you know, you just like getting these aha moments, just like...
really justifies like our approach here.
I think like by writing many, many of these like domain specific skills or task specific skills, we're getting a lot of value and just, you don't even, in this case, like, yeah, like I said, I didn't even know that this skill existed, but it did the right thing and figured it out.
And that's what we want to see across all work.
Even at Intercom, like AI adoption is unevenly distributed.
We've got teams, people, you know, all at different levels of maturity.
And we do a lot of work, enablement work, and that to help them understand or give them space to do things.
But I think Steve Yeagy recently released a...
or talked about maturity rating for engineers about like how AI pilled somebody is.
And I think like this, this, this is like how I think about this internally inside of Intercom.
It's similar enough, but kind of goes in a little bit of a different direction.
But, you know, and, you know, depending on the task and stuff, sometimes I would like regress.
Sometimes I'm, I'm at step three or four or something.
But the first few steps are definitely like, you know, people just adopting, getting used to these things and then starting to.
only use the tools to produce code.
But then later, the way we want people to work with this stuff is use Cloud Code for everything, automate your work, write skills, write really good skills, write skills that update the skills.
And then you start to optimize the environment for the agents.
and we're starting to see this in our software architecture in uh the different decisions that we're making that we're like we're starting to bend or kind of shape our environment so that uh agents just have an easier time or you can get more done um and so an example of this would be uh we've dealt with um areas where you might be plugging in a load of different say third parties into us like a say messaging uh providers uh and they all have kind of similar APIs, but they can kind of differ in implementation and details.
And so we've had a lot of success by building like these really great solid cores and then being able to rapidly kind of let agents like do all of the annoying work of looking up docs and integrating with SDKs and all of that stuff at scale and plug into this like solid core.
And I think it's probably good software architecture anyway, but it allows the agents just to move super fast here and get like really great, really, really consistent output.
So here's where we're at.
So actually, we met this goal in the last few weeks.
So we doubled PR throughput in nine months.
And we actually tripled PR throughput in 16 months, which was, we only kind of recently realized this.
You can see a wild inflection point as well around December 2025.
Again, we had like the agents getting better, harnesses getting better.
But that was also the same time that we decided we're going all in on one tool.
We are, like we had built out the full-time team to support and setting up skills and setting up all of the things to make things really, really easy for people to adopt.
So today, yeah, we are, pure throughput is well over double what it was nine months ago.
And we see no sign of this stopping as well.
This is continuing to go in this direction.
Some data I pulled out last week.
So we're at 95.9% of...
pull requests being authored by Claude.
We also have the bottleneck of approvals and so we've been doing something about that.
We now have Claude approved, like fully identically approved pull requests.
At the moment it's around 17%.
We basically work with Claude to define what's extremely safe and we're confident that these pull requests can go to production without.
without any additional oversight or without having to kind of force a human to be on the loop on it.
And we expect as well, like we're going to, we're continuing to work on this.
We want over 50% of our pull requests to be fully approved automatically.
And we've been doing work, you know, working with our auditors, making sure that there's no risk here to SOC 2, ISO 27001, et cetera.
But the thing about these approvals, I think they're actually being done better, like a higher standard than humans would have done.
They're consistent.
They never forget anything.
And I think as well, just kind of like water flowing down a hill or something, we think the work will kind of...
bends towards the path of least resistance.
And so when people kind of see the kind of shape of changes that they can make and get automatic approvals, they'll kind of shape all of their work towards it.
You know, the pull requests have to be small, well, in terms of code changes, feature flags used whenever possible, metrics available, observability things.
Like basically, they just have to follow all of our best practice that we would consider to be good.
And the description of the change actually.
matches what the code kind of does.
And so we expect this number to increase a lot over the next while.
Skill invocations.
So we've been writing hundreds of skills, lots of people using them.
We kind of measure these things a few different ways.
We send all metadata bit sessions to Honeycomb.
So we can see exactly who's calling what.
different bits of information about individual cloud code sessions.
And this is like internally available.
You know, I can go into these dashboards and look at them.
But we also pull out session transcripts.
So we hook into cloud code, use hooks and things like that to copy the session transcripts for every single session into an S3 bucket.
We anonymize them and we then data mine these for kind of insights.
Or they're just very handy for tech support as well.
We have hundreds of people using these tools.
When something goes wrong or something goes awry, we want to know about it.
And it's really, really useful to have the session information there to hand so that we can proactively go after or improve things by looking directly at the session data and not just relying on the human to tell us what went wrong.
This thing, this graph is like, maybe one of the most interesting ones.
This is our defect list.
I'm not particularly proud of this ever-growing defect list.
This is over the course of, I guess, a year or so.
And we haven't put a huge amount of focus on this.
This hasn't been a goal or a target.
But teams are getting through more work and they're getting through more defects.
And some teams have been pushing for things like defect zero.
And so there was a question asked of another talk, like, what do people do with their time?
that they've gotten back by using these tools.
And one thing that they're doing in Intercon is closing defects.
And so we're actually down now like over 50%.
So I think this is a little out of date.
This is like a month old, which is wildly old.
And we're now down at more than, like we've removed more than 50% of defects from the peak whenever it was in January or something like that.
We've also been looking at some other data.
Like we're worrying about, say, code quality.
And we've been...
partnering with a research group in Stanford, we shift and give them all our code and they kind of give us some metrics.
And over the course of this year, code quality did kind of start to go down a bit based on rework and complexity and a few different metrics that are looking at.
But then over the last few months or weeks, it's gone in a completely opposite direction.
The average code quality is now, or the average additions that we're adding to our codebases, they're improving the overall quality of the codebase.
Again, it wasn't a specific target.
Definitely something we're interested in.
But it was amazing that this kind of just got turned around by us working with the agents, giving better guidance, linting, all of the kind of guardrails.
And we're seeing this, like, objectively we have data from researchers that shows that things are going in the right direction which is really interesting um other stuff that we've been looking at as well has been like time from initial code being written to the time when a feature is announced so that um and like that's been compressing as well which is interesting um and yeah this this data is like these are These are like training.
These are not targets.
These are just things that are happening in our environment.
And I'm very interested in like maybe pushing teams to go after more teams to go after Dpx0.
So going through like our cloud code setup, we have like hundreds of plugins.
We have dozens of plugins.
There's like 42 plugins and hundreds of skills.
and we've been breaking down access and giving access to as many things as we can.
So anything I can access my laptop, the agents must be able to access as well.
And so that means like fully production, full data, everything.
And the plugins, so we have these like, we have core plugins, plugins that like are there to like make sure the cloud code is working exactly the way we want it to.
And we have extremely high quality skills that are used often by like say every single software engineer.
And we take time to make sure that these are high quality and running valves against them and really push them on the quality side.
But we also want it to be easy for people to distribute.
skills, especially on teams and things like that.
A lot of the work is just like local to teams.
And we can have like a more relaxed quality bar when it comes to non-core things as well.
You know, we want, we want, we don't want to be.
gatekeeping too aggressively here.
We want people to try things out and then maybe if skills get adoption or we can use the session data or telemetry data to kind of see what's being used and then where to invest and like find gaps and issues with them.
So yeah, we have like hundreds of people contributing to these things, thousands of changes going through, loads and loads of test files.
But the main thing is the we we have an extremely high quality bar with any kind of individual skills skills need to be small composable testable um and uh like not just trying to do a bunch of open-ended things, but, and we really, really push for that quality bar on, especially our core skills.
You know, no one wants to, you can't build a senior engineer out of a bunch of skills and tasks that like are only say 50, good 50% of the time.
It's like, that would not be a high performer on your team.
And so really sweating the detail, sweating on the details and quality.
um is like absolutely important like critical for uh widespread successful adoption of these things um so like just looking through some of our core skills this is like these base plugins like uh so everybody gets these on their laptops and force we force install these things uh we actually bypass the cloud code kind of update mechanisms and stuff for getting these skills and things onto people's laptops we use our it systems to publish uh the all of the plugins directly down to people's laptops as well as cloud code configuration and other things this bypasses all sorts of issues with cloud code and makes sure that we have we know exactly what people are running And then, yeah, the base plugins run absolutely everywhere, does all the telemetry work, some safety things, just generally making sure that we know how well set up and the configuration is in place.
And then we just have dozens and dozens of specs or skills that do individual tasks and do it well.
Fixing flaky specs, these are like fixing flaky tests.
It's a skill I wrote, which does a really, really world-class job on like...
on fixing tests that are intermittently failing.
We have hundreds of thousands of tests.
We run our test reads thousands of times a day.
And so you just end up with a lot of these kind of tests.
And because we're shipping so often, it doesn't block shipping, but it kind of slows things down.
It's a bit annoying.
We care about these things.
We open issues for them, but we don't.
And we aggressively skip them as well.
The value of any individual test is actually pretty low.
But this is a skill that I pretty much iterated with Claude Code to generate in a feedback loop.
Just got it to fix dozens and then hundreds of flaky specs.
And so I didn't write all this down up front.
I didn't design this whole skill or whatever.
I just iterated with the agent and iterated, gave it a goal and gave it lots and lots of data.
We have hundreds, thousands of these issues from historical data.
And then you end up with this extremely detailed step-by-step, like...
play by play here's all of the different individual steps you have to do to like do a world-class job at like recreating an issue finding what what it is and then also you end up with like classifications and so like these like classifications these are like total cheat codes and it's like how we work as well if I join an incident um I'm the first thing I'm thinking is not like okay I'm gonna start from first principles uh work out what's going on i kind of just go does this look like a database issue uh it kind of smells like like it is um and you get to faster outcomes with these kind of like cheat codes um and again this is interconceptive context these classifications are like most of them are kind of universal for any rails app but like they're specific to our environment and the skill as well has got it in it like if you learn a new uh classification if you new something new kind of comes up update the skill and so the skill is kind of self-updating in that way and now just like can resolve basically any flaky test um really really quickly really really accurately and uh like at the standard of like way way better than i could do uh at the standard of like our best rails engineers other stuff that we're doing so i've been talking mostly about like the R&D org here.
Cloud code's kind of gone viral across all of Intercom.
We have over 1,000 weekly users.
Something like 80% of people in the Intercom are using Cloud code every week.
And yeah, just like democratizing access to data.
And people are just doing way more analysis, getting way more information about customers or things.
Using Cloud Code, it's been absolutely going wild and viral beyond engineering.
Other things we're doing, like say right now, yeah, replacing all Rumbux.
We don't want humans acting like troubleshooting outages or dealing with alarms.
It must be agent first.
We're working on remote agents, kind of similar to the Airbnb we're kind of building.
We want to move this work away from people's laptops and get next levels of throughput and scalability through that.
And yeah, like everyone else, we're also worrying or thinking about what does this mean for our jobs and roles and all of this?
Things are just merging.
But I think it's too early to make any kind of big decisions.
But we've been running experiments and trying things out to see what this new world looks like.
And certainly, I think the planning, team organization, everything, all going to get changed over the next year.
Anything else?
Oh, yeah.
I've been working as well on actually shipping products, shipping features.
And it's amazing.
I've just been using Cloud Code skills to do all of the product management and all of that stuff.
I wouldn't have gone near this work ever in the past before.
But now I'm able to ship stuff to the internet.
And so this is not directly related to the AI productivity stuff.
But it's just me personally, I was able to move my role to be more like doing product management stuff and this, that and the other and released a cool CLI for Intercom.
So that's the talk.
I wish you all the best of luck.
If you aren't doing pretty much all of these things today, you're going to be doing so in the near future.
We published an update.
like a really big first part of a blog post, which is a series on our use of like our 2x project.
So there's more detail about the stuff here.
It's on ideas.fin.ai.
Fin.ai slash CLI is where my little CLI is.
And I've got like links to my talks and stuff like that on brian.scanlon.ai.
Yeah.
I'm around for the rest of the day as well.
More than happy to have chats and yeah, more than happy to answer questions as well for the next few minutes.
Okay.
So good, Brian.
The chat is blowing up.
People have lots of questions for you.
When I was off stage, I had asked you, were there questions that you were hoping got asked and you offered to tell us about your costs.
So I know people in this audience want to know how much this is costing you in token spend and everything.
So what does that look like?
Yeah, we had, I think two weeks ago was our highest week so far.
Like, so like everyone else, our costs are like that.
And we had a nice round number.
It was like $128,000 a week.
And yeah, but we had like two quiet weeks because of the Easter break over the last while.
So it's kind of sort of stabilized.
But like, I'm pretty sure like next week it'll be 150K and then next month it'll be 200K.
Like, this is going in one direction.
150k per week and you said you have roughly a thousand developers?
No, we have a thousand people using cloud code.
We've got about three, four hundred developers.
Got it, got it.
And some developers are spending 20k a week or a month.
Yeah, it's adding up.
And you don't have any caps in place of like, hey, you can't, once you get to this spend, you're cut off.
It's just use it, do what you can.
Yeah, we have infinite spends.
It helps that we have our finance team using Cloud Code.
So they're like, oh, wow, this is cool.
And one of the interesting things we published today on our blog was the cost per pull request has gone down.
And that's the way to think about this stuff.
It's not just like, oh, this cost, and this cost is annoying.
I think you should be thinking about it as a cost that's additional to human salary.
Yeah, we're getting through more work cheaper than before.
Of course, it's a big bill, and we do want to optimize it.
And we are going to be doing some stuff to go deeper into optimizing our spend.
But right now, I think the opportunity cost is greater than just get out of people's ways, let them do whatever, and pay the bill.
Pay the bill.
You heard it here first.
Okay, well, some of the other questions that are coming through is what happens if some of these tools change their pricing model?
What happens when the subsidiaries go away?
What happens if they 10x their cost next week?
Yeah, I think there's real business continuity risk in the way, like how aggressive we've been adopting the tooling.
So Stephen, uptime is an issue, I think.
No one has used these tools recently.
That hasn't run into some outage over the last while.
And we're doing basic stuff there.
We can failover provider.
We can move to Bedrock, AWS-hosted models.
And at some stage, I think going multi-provider is reasonable to think about from a business continuity perspective in the kind of like, yeah, what happens if Antropic falls off the internet?
Or what would it take to kind of move over?
And then I think, of course, as well, the open source models are probably OK.
I've got to assume they're probably going to be OK enough in the next while to be at the standard where Cloud Code is today.
So if things go down a not very nice path in terms of cost or availability, I'm pretty convinced that the open source models will be good enough to at least do the basic functions.
So that's like insurance policy, I guess.
Always good to have an insurance policy for sure, especially in today's age.
Tell us a little bit more about how you are stress testing these skills.
How do you audit them?
How do you make sure that they're up to date?
How do you make sure that a junior developer who just started at the company doesn't create a skill that everyone starts using and is completely wrong?
Yeah, paying a lot of attention to the quality and outputs is important.
We've had that exact situation of where, say, less experienced people trying to automate some work.
And my favorite story is we had some, I think it was some interns that were looking at resolving some exceptions, like errors in our application.
And Cloud Code is very literal.
It'll do what you tell us to do.
And they kind of told us, like, hey, fix this exception.
And the exception was coming out of a...
a custom emoji in a discord message and it was just kind of blowing up um on the custom emoji and clog codes like fixed the out of it by like dropping the message if it had a custom emoji and it sounds like that's not really what you should you know um and yeah it's it's so it's a bit of a silly story but the point is that like you really have to understand like the uh like be able to judge the output of us and uh so we like we've socialized that as an issue like internally and we talk about the quality aspect for a lot um but also we don't want to gatekeep we want people to to yeah do like we do want our interns or whoever uh to be knocking out some skills getting some usage out of us because like we're on a journey and you have to learn you have to get some stuff in and use um and so that's why i like things like our plugin setup where people can publish these things say locally to their team or whatever and you know that kind of minimizes the damage or potential but then When people want to do a better job, they notice maybe it's not working right.
We're there to help.
So we can help guide them around what good skills look like.
And we have skills that grade skills and probably a skill grading, that skill grader.
And so you can get from a half working skill to something that works really well very fast.
And we aggressively use evals and tests to make sure that they're doing the right thing.
We have skills for skills.
We have AI for AI.
Got all the things.
Okay, maybe one last question before we wrap.
Something that we've heard in a couple of talks today is this concept of being able to use agents and AI to help auto-approve PRs.
You mentioned that you all have about a 50% response rate, and that's still a work in progress to make that even higher.
What are the criteria of what gets to be auto-approved versus what doesn't?
Yeah, so the pull requests have to be small.
I think it's pretty strict at the moment.
It might even be as low as 50 lines or 20 lines or something like that.
So we just simply refuse to automatically approve anything that's particularly large.
And if something doesn't have tests, we're not going to approve it.
We do have sensitive code areas.
There's definitely different parts.
We kind of give high-level guidance to Cloud Code and say, yeah, if it's in a high-volume code path or it's a critical database transaction or something like that, just don't auto-approve it.
That's kind of it.
It's not that sophisticated, but we built the process to develop the criteria by analyzing hundreds of thousands of actual pull requests, picked out what good looks like, and then we got humans to grade it.
So we got it to go through some sample PRs.
it would get the results and then we had like a human like labeling these results to see what the uh the like percentage rate is like is it getting like 90 or 50 or whatever of like that that it would make the same decision that the human does um and once we got into like the upper 90s which was pretty quick we're like okay this is ready to go um and we're looking at like increasing the uh the criteria a little bit, like loosing them up.
But really, it's like we just have an opinion now about what a great pull request looks like, and we want to shape all work towards that.
And so rather than being expansive under what we allow, we want the work, the inputs to be changed.
And so I think that will get us the biggest results over time.
Amazing.
Thank you so much, Brian.
Thanks for traveling all the way from across the pond.
Of course.
Thank you.
