# AI-Driven Engineering Velocity and Quality Guardrails

**Podcast:** Engineering with AI
**Published:** 2026-06-15

## Transcript

Oh, it's going to be controversial.
I think pair programming as we're used to it, where one person is driving, writing commands, the other one's looking at the shoulder, making sure it's correct and learning from it.
That way kind of broke when you started using multi-agents, right?
However, very important to have senior engineers sit with more junior engineers to show.
how you're working with AI, right?
You just assume that everyone knows all these tricks I got off my sleeves and also like writing the AI workflows together, I think is also very important because they're so critical.
But the traditional pair programming paradigm where you just sit and watch over each other's shoulder, I think that one needs to be replaced with a different flavor of it at least.
I don't think that's getting side by side for eight hours a day is that productive anymore.
Jonas Clayson is a principal engineer at Provision Analytics, which makes cloud software to streamline food safety, quality assurance.
He's also a Navalya alumni and a ThoughtWorks alumni.
He and I worked together at Navalya, actually, not that long ago.
And recently, he's been posting stuff on LinkedIn that I really love, the sort of concrete stuff that just makes me smile because it isn't just sort of platitudes or predictions or forecasts.
It's like, no, this is directly how AI is changing what I'm doing.
talking about wasted time providing SQL fixes and how work trees are better than branches.
And probably my favorite was him talking about having his laptop next to him in the car and running cloud code as he was driving home.
So the content should be refactoring.
I am just loving all of this stuff.
And I've been working on getting on the podcast.
Today, Jonas is our guest on the Engineering with AI podcast.
So Jonas, thank you so much for being on the show today.
We really appreciate it.
Thank you so much.
I'm very excited to be on.
I've been with these episodes all along.
So it's really great to be on here.
Amazing.
Amazing.
No, I mean, it's funny.
I'm used to talking to you in this exact situation.
This is my home office and that's your office.
Yeah, I love seeing you drop your basketball every time.
Well, and now that this is my studio, it's literally on there every time.
And it's too bad about last night, but I don't know.
We'll see.
Maybe.
Game seven.
Yeah, game seven.
Hopefully BI will come back and we'll be okay.
But anyhow, let's get past that and into the software development.
Like I said in the bio and in the intro, I've been loving reading what you've been talking about.
You know, so much really good stuff.
We're going to link to some of it in the show notes.
What's one of your favorites?
What's the one that stands out for you?
Of the articles?
I think they kind of follow a similar theme.
It's like how AI, everyone talks about AI just makes cogeneration cheaper, right?
But it also makes engineering rigor cheaper.
There's so many things that you can now do to make sure the quality of your software is better, right?
That you could still do before, but most of the time didn't have time for it.
It was just too much effort.
But now you can have AI do all these extra checks and tests and validations and everything, right?
So I think that's one thing a lot of people are not talking about.
Love that.
And this is a really good time for you, too.
I mean, you've been at Provision now, it feels like just yesterday, but it's actually been a while.
Tell me about what's going on there.
Like, maybe start with a little bit about the business and your role.
But then, like, I'm guessing this has been timely for you guys.
So let's start at the top with Provision.
Yeah, so it's a big SaaS company that are just in the process of rewriting their software from scratch, make it more...
scalable and easier to use and everything, right?
So I came in a really good time where we're just starting this process.
And we were starting with CoPilot, right?
About a year ago, that was kind of the only game in town.
Claude was slowly coming along.
And it's been just interesting to be able to experiment when you're in a small company and you have a CEO who's really bought into AI and is very supportive.
We've been able to just go all in on this.
I think I mentioned earlier to you, I think it was November, December last year, I just made the decision.
I'm not going to write a single line of code anymore.
I'm just going to go on with this and Claude's going to write all the code for me.
And then the whole team kind of got off to this.
So now we have an engineering team of six developers that are just using agents to write code.
Wow, cool.
Okay, so briefly, I mean, not in an exhaustive detail, but like what's the stack, what language and that kind of thing?
Yeah, so it's...
front-end react test that query kind of what most people do uh back-end is dotnet with the postgres database a couple of other external services for sending out emails and notifications cool it's clerk for off um so pretty standard simple not simple it's got a lot of complexity but um you're also my first guest who's using claude with c-sharp i i suspected it would be good i've had it look at like one c-sharp code base so far But, but yeah, I'll be interested to get your take on that.
That's cool.
Yeah, that's one thing I noticed that the more established technologies to use, the better the AI is.
If you use some kind of weird library, there isn't that much code out there for the AI to kind of learn from.
So if you use these established languages, it does a really good job.
Yeah, for sure.
Okay, cool.
Okay, so, so, so great.
We sort of, we sort of already have to do it.
Let's start with you.
We'll get to the team and all that.
But let's start with you.
And I think you've already said it.
Like you're saying you're having Claude write all of the code now.
So I guess I know the answer.
What's your favorite tool so far?
I mean, I use Codex as well once in a while, but Claude really fits my workflow.
And especially I use the CLI and having multiple tabs and just quickly go between them and check on your agents and all the shortcuts and everything.
It just works really well for me.
But I'm very tool agnostic.
I constantly try out different tools.
Try not to be too married to one of them.
Yeah.
I wish I did.
I have like a productivity quirk about me where I don't like, you know, that comic where the guys are like, hey, look, it's a wheel.
And the other guys have like a wagon that has square wheels.
I'm the guy with the square wheels.
But I make sure that I have smart friends around who will tell me on occasion, like there really is a better wheel now, Kyle.
But personally, I'm terrible at actually making myself do it.
So, yeah, cool.
So I'm.
I'm glad you've tried Kovex and those kinds too.
I'm with you.
But like you say, like plugs, probably one of the better ones right now.
That's still holding true.
I think so.
Yeah.
But yeah, changing so fast, you got to stand up.
Well, right.
Well, and you know, that's a question I've got.
Somebody once said like, oh, but you're switching costs.
I don't know.
Like if I were a Fortune 500.
right, legacy business.
And I'd spent a lot of time with the salesperson, you know, at Anthropic getting the master services agreement just right and the statement of work just right and making sure that like, you know, all the language I wanted in there about my intellectual property, maybe I have high switching costs around the contract, but are there any other switching costs?
Like I've had like code.
Well, actually it wasn't code.
It was Gemini CLI.
Check out a code base that I've been writing with Claude in it.
it just read the CloudMD.
Like these are all text files, right?
So it feels like switching costs are going to be really low in the end.
How's that feel to you?
Is that right?
Yeah, definitely.
I can't even, I mean, we got a lot of custom Cloud workflows, but they could easily be reported.
There's just some things how a Cloud deals with work trees and how they enter different work trees when you start them.
Sure, we've got to figure out how to do that in Codex, but.
In general, it would be fairly easy to switch to anything else.
And we're actually using Codex as well.
And we'll come to that later when we talk about the PR review, because we use both Claude and Codex.
And then we have them battle each other.
Nice.
Okay.
In one of your other blog posts, you talked about using the Ralph Wiggum loop.
And this reminds me, I played around with Gastown, which is...
kind of similar.
It's different, but it's kind of trying to get you to that like unattended operation.
And I think that's what, you know, the Ralph Wiggum was trying to be all about too, is give it some exit criteria and let it just turn.
So that's cool that you've been playing with that.
A lot of people are, you know, reading about it, but not necessarily doing it.
Yeah, I was just waiting for an example where I could apply it.
To be honest, I had a perfect scenario for it.
Tell me about it.
Yeah.
fit the sweet spot of when to use it so we have all these integration tests for all our api endpoints hundreds and hundreds of them right and they all insert c data calling the api in certain database and then calling these endpoints to like list and filter after they've inserted data obviously not very optimized right because you have to do a lot of inserts and not that great because you won't have that much data in the database either for your your all your list operations and filtering and stuff like that so the ob is solution to this is just, hey, let's just load a database with some read-only seed data for one organization.
And then anytime you need to test like a get operation, you just work against the seed organization, like the read-only data.
Before AI, that would have been sitting on the tech debt column a long time.
No one wanted to touch that.
That's like a week of misery of someone just figuring out every test.
Okay, I need to move this into this seed data as read-only.
So what the Ralph Wiggum technique does, it basically takes one test at a time and it starts with a fresh context.
So you're not going to start hallucinating because this was running for hours.
So we just take one test and then keep a list of all the tests it had to do, right?
And every time it runs in this loop, picks the next test that it hasn't processed and figure out how we can add more C data to satisfy this test.
So it didn't have to do all these inserts, right?
So it probably cut off like a 30.
40 seconds of every PR that ran all these integration tests.
Wait, wait, wait, wait, wait.
Okay.
So I followed a lot of that.
The Ralflake Wigam loop had like a stack of tests that it had to go and figure out what seed data to create.
So instead of doing these like manual inserts as part of the test, it just had like in the database already.
Like it was just a big insert in the beginning and then all the tests reused that data.
Oh, so this was about accelerating your, like your continuous delivery build times.
Yeah, just refactoring all these tests, which is such an unpleasant thing to do manually.
But with AI and Ralph Wiggum, that's another thing that's very important, this feedback loop, by telling it, like, when is the AI done?
Just don't tell it, like, hey, go do this.
What's your definition done?
How do you verify yourself that you're done?
So how much time did it shave off the build?
Maybe 30, 40 seconds off, like, at the time, two-minute build.
So it's pretty significant.
Holy cow.
But that like a big percentage wise decreased.
Yeah.
And I wrote the prompt in like less than an hour.
And then it just, that was a Friday evening.
It just, yeah.
Ralph just kept going and going.
And I can't imagine like, you know, facing down a product manager or engineering manager and try to say like, hey, we've got this thing we want to do.
It's going to take a week, but it's going to save 30 seconds per bill.
Yeah.
Yeah.
That's a hard sell.
Never going to get greenlit.
But now that it's down to an hour, you don't even ask.
You just do it.
Yeah.
So, yeah.
That's one thing I noticed, like our tech debt is so small because you don't have those things that you don't want to touch.
Like anything is pretty doable now.
Right, right, right, right, right.
Well, I think that's smart.
What you're inferring there is that you use the speed, the time that you get back from the speed, you're not using it to just make more product faster.
You're using it to pay down tech debt too, right?
So I think that's an important part of what you're saying there.
That's where people get it wrong and it's very tempting.
What used to take four days, right?
If it takes a day now, you freed up three days.
It's very tempting.
Just like, I'll just grab some more stores and just get more functionality out.
Yeah.
That's what gets you in trouble.
You have to take that time to invest in the product.
Yeah.
Take that, write more tests.
You got any way that you can just make things more stable and smooth in your deployment pipelines or wherever you need, right?
Just you have all this extra time.
Use it wisely.
Yeah.
Okay.
So one of the questions I ask everybody is what card rails are you putting in place?
And you kind of hinted at something around codecs and PRs.
Maybe this makes sense to tell the story now.
Yeah.
I hear a lot of people complain about agents doing PR reviews.
And to some extent, they're correct.
Like if you use like GitHub Copilot and it's reviewing any PR in Git for you.
We always joke in the team.
It's almost like...
GitHub Copa is going, oh, I found this issue.
Oh, I found this issue here.
It just tries to be so helpful.
And like every little issue is like top priority, right?
It doesn't have this broad context and really narrowed down on like, these are the things you really need to fix.
These are a lot of nitpicky things.
So we have this workflow now where every developer runs a custom cloud command called PR review, where we have a very opinionated.
I think this is the key.
You need to write a prompt.
It's very opinionated.
This is the things we care about.
Don't just give it like, Claude, you figure out the review.
This is what we care about.
We care about load complexity.
Don't duplicate code.
Don't try to introduce new frameworks or patterns.
Keep it simple.
Look at the rest of the code base.
Be consistent, right?
So it's this prompt that is super opinionated on how you build software.
So we run that burst.
It always categorizes like a must fix, should fix, or like nitpick stuff, right?
So developer runs that.
The entire team runs it.
And you go ahead and you fix those issues.
Then you could open a PR in GitHub and then we have a GitHub action that calls Codex and it kind of does the same thing, but a little bit different.
So Codex runs the same thing then, right?
And surprisingly enough, it finds stuff that called misses.
It's more for bugs and stuff like that.
And I can imagine that you also put this in your economy to avoid these kinds of things.
Yeah.
So it missed it the first time and it missed it the second time in the Pulver Crutch review.
Yeah.
And then we have another command, right, that pulls the Claude comments and the Codex comments.
And it kind of gives this unified view.
And then also super opinion.
It tells it, okay, be honest now, Claude.
Go for your comments.
Go for Codex comments.
Which one should we fix?
And it's surprisingly good.
And I usually agree with it.
Wow.
So it's, okay, so it's reading its own work.
It's reading Codex's work.
And then it's coming to you with a final opinion, which you then validate and say, like, yes.
Okay.
Yeah.
I know it might happen.
interaction with it.
Okay.
Tell me more about this one.
What do you think?
And then I tell you, okay, go and fix this.
So then we have two rounds of code reviews by AI.
Then we assign a human because that's always, that's the bottleneck right now.
It's the human review.
So we're trying to just make that as smooth as possible.
So it's more of a quick sanity check, right?
If I look at a PR, I might look at which data migrations, which new endpoints, like stuff that matters.
I'm not going into the things and like, oh, you used a for loop there instead of a for each.
Claude already took care of that, right?
So I'm more laser focused on certain things.
So it sounds like a fair amount of the guardrails here are getting rid of false positives, like trying to prevent it from telling you stuff that doesn't matter.
How good is it at finding stuff that does matter?
I've been pretty impressed by it so far, like way more than like we used Copilot originally.
Like we still had that clutter view and then it went to copilot and copilot would just find all these like little things that didn't really matter.
It was just so much noise, very little signal.
But yeah, it's always about tuning these context paths as well, right?
Like this, this prompts, if there's things that's missing or if you think it's focusing too much on one thing, we'll go and tweak it a little bit to make sure it aligns with what we value as good quality code.
Yeah.
Okay.
But it's still finding stuff, right?
If it weren't finding anything, you would have like turned it off by now.
All the time.
and ship a PR without cloud finding stuff and codecs usually finds one or two items as well.
Yeah.
And I know there's a huge difference.
Like everyone on a team, when I look at their PR, even if they're more junior engineer, like before it would be like a learning experience, like you find things and try to teach them.
It's a whole back and forth.
Now I look at a more junior PR and it's solid.
It looks like a senior engineer wrote it because they've been guided by all these context files and AI automation workflow that they're using, right, to make sure that their code is aligned with exactly how we want it written.
Cool.
Cool.
Okay.
You're talking me into it.
I'm actually at the point where I'm like, I just don't do them at all.
I'm not doing PR reviews.
I mean, for me, it's a one-person show and it's a pre-revenue startup.
So that's an easy decision to make.
You know, this is a side gig.
It's not even my day job.
We're obviously not doing that in enterprise.
In enterprise, there's still PRs going, right?
So I'm thinking about teams that I know at work, right, that could use this.
And actually, you know the team I'm thinking of too.
So one of the developers on this team developed a similar kind of bot that is all aimed at Sneak vulnerabilities.
So it goes through whatever Sneak is reporting and it like finds any vulnerability and fixes it.
And one of the things that they sort of tacked on is they made sure to run their end-to-end test, which you know the EDE test I'm talking about.
It was kind of broken like all of the time.
now that we have Claude, it isn't, you know, and so it's actually kind of useful.
And so it's been a nice tie-in to have it, like, reduces the effort with keeping us up to date in Sneak Home capabilities, which is kind of cool.
But we don't have your PR wizardry yet, so I'm going to tell them about that.
Awesome.
I play with it.
I've been surprisingly impressed with it because you think, like, you have Claude write all this code, and then at the end of it, you tell them, okay, now go review it.
Like, it already told you, like, it's not going to find anything, but it does because it kind of...
step back or to get a new prompt.
It looks at things from a different perspective and it finds stuff all the time.
And a lot of times it'll find bugs that I would have never found, right?
Like these small little things that are very easy to miss when you manually look at it.
Yeah.
I think superpowers, which I'm using right now, one of the cloud plugins, is it is finding a few things.
I think it does a little bit of what you're doing, but it's not nearly as elaborate as what you're describing.
Okay, fine.
I'm convinced I'm going to try it and I'm going to tell the teams to try it too.
Okay, so let's get back to you guys.
So we talked a little bit about you and we've talked about some of the experiences you guys have had.
Let's get into the teams that you're serving, right?
You've probably been there for the arc here.
How has product management and story breakdown changed?
When you think about how provision works, like, do you have product managers that write the PRDs and then with Claude and then Claude takes the PRDs and turns them into stories?
Like what?
Probably not, right?
Like, so, but what has changed with both product management and story breakdown and things like that?
The whole company uses cloud across the boards, even product, customer service, everything.
So it's very ingrained in the company.
One thing I've been thinking about related to story refinement, which we don't do right now, the first thing we do when we start a story is we have like a custom cloud command, start PR, you give it the ticket number, it goes to linear.
It grabs the whole description, any figment designs, any attachment, anything it needs, right?
And it starts querying you about gaps.
Like if things are not in the store because it has access to the story and the code base, it knows like, well, you didn't tell me about these three items.
And that's the developer doing that right now.
I will tell it the gaps in the technology, but if there's like business requirements that are missing, it will kind of tell me about that.
And I have to go back to the product owner.
It was usually super fast and will tell me right away when I when I slack him.
But it'd be cool to move that up before the developer touches.
So the product owner could run some cloud command that has access to code base and help him refine it.
Like tell him, well, developer can't start working on this because you haven't.
define what happens in these scenarios, right?
If you archive this employee here in the system and then you go here and should it show up in the dropdown or not?
Like all these things that could be refined before it even hits the developer.
Right now we get it on the developer.
It's not a huge thing, but it would be nice if it was just like completely refined.
We pick it up and Klaudah's already filled in the gaps before a developer even touches it.
That's fantastic.
Yeah.
Honestly, that's...
the kind of thing that I've been like waiting for and one of the first people that's actually called something like that, you know, to have it as glue between us.
Yeah, I'm probably going to work on a half this to be honest.
So I'm like a refined ticket.
Just give it a story and it will basically go and update your linear based on your answers when it's kind of telling you what's missing.
Nice.
What about architecture and planning?
To be honest, in a startup, I mean, it helps a lot with just.
prototype in different architectures and validate it in there because it's so quick to just spin up different solutions in architecture.
Working for a smaller company, it's not as big of a challenge as working for a big enterprise where you have a gazillion microservices all over the place and different things, right?
Like we're a pretty lean team.
But we recently had to figure out how to do analytics on our product.
And it was very good to just kind of like, hey, we'll meet all the ways in which you can do this.
We actually have a very complicated analytics problem.
Because all our data for all our customers, all our customers, they can create their own forms and they can create whatever fields and grids and tables they want, right?
So nothing is predefined and they can have multiple versions of that.
And every customer is different.
And then we need to get that into a warehouse and expose that as like a flattened out structure.
We actually met with our Microsoft representative.
He was going to demo Fabric and he was like, oh, look at how good this handles JSON.
And we told him the challenges we were facing and he like almost felt like he was starting to cry.
Like, God, no idea how to solve this.
But with Claude, you can tell it is problem.
It's like it practically told it there's no product out there that can really solve this for you.
But here's a couple of ways you can do it.
And we actually came up with a solution fairly quickly that was able to like extract all these.
completely unstructured JSON are different for every client into like a unified data warehouse that is being updated depending on the structure and the JSON and the schema.
Wow.
Cool.
Did you end up using AI agents as part of that pipeline?
More just to come up with the solution and to build like the custom code to make this happen, which would have been a lot of effort.
It was also really good to have AI there because it's very easy to test with AI.
You can just tell you built all these new views in the database that are like the flat nut structure.
You just go and compare every single view with the original JSON document.
I know it's going to, that would be horrible to do manually.
Yeah, yeah, yeah.
This is the kind of thing that software developers would have fought about for, well not fought, debated about.
for months endlessly about like, wow, I can't do it that way.
It'll never work because I read the book once and, you know, and somebody else will have, oh, I saw something on in a magazine about such and such and well, it wouldn't be a magazine, but you get the idea.
It's the kind of thing that people would have like debated endlessly and never started because we couldn't agree on how.
But to your point, you can just sort of like wade in and say like, what do you think?
Just go.
It's not a lot of cost.
Yeah, just the cost of prototyping.
Just do it this way.
Let's try it.
Let's see how that works.
Let's try it this way.
You can do multiple prototypes.
You just didn't have that luxury before AI, to be honest.
Right, right, right, right.
100%.
And just all the QA can do for you, right?
Just like, just spend the week and just go and like do extensive testing on all these things I built.
Like, chug through the data.
What other changes are you seeing in software development?
I mean, this is kind of like all about software development, I know.
But like, at the actual like dev time.
I mean, it's really changed how you work, right?
Like, it's...
you're an AI coordinator now.
You're not a software engineer anymore.
You're more guiding your agents and having them do the work, reviewing their work, inching their context file to make them work exactly what you want.
Yeah, so that's essentially what we've been doing with.
Right.
Well, it's changed the role, to your point.
Yeah, absolutely.
Yeah, I think the rigor in the engineering now isn't about the four loop or the four-inch loop or the wild loop.
It's about this.
It's about how do you make sure that it doesn't make that mistake again.
Yeah.
Okay.
What have you seen in testing?
How does it change how testing works?
Oh, that's a big one as well.
Let's start from just locally when you're testing, right?
It's so important to tell AI what is considered done.
Like I mentioned before, if I write a new endpoint in API, I always tell, okay, don't just write the code.
This is essentially the integration test I want you to write.
And it's going to send this data.
It's going to do these assertions when it come back and just keep working till this passes, right?
And like, don't just tell me go right and then I'll test it.
Like you have to tell it when to stop.
And the same thing on the UI side.
I do this all the time.
It's like I could get AI to write some code and then open the browser and click around, make sure it works.
But I always tell it like, hey, launch the Playwright MCP.
You already know the story.
You have all the context.
You know what you have to test.
Just go and do it.
People think, oh, well, it's really slow, right?
When you run the Playwright MCP.
see through cloud.
It's kind of slow because it needs to navigate.
It needs to figure out which buttons to click, but it saves a lot of time because you don't have to constantly tell, oh, it's not working.
I'm not seeing the button, right?
When cloud is navigating the MTCP and it's going through the DOM, it knows it has all that context so you can quickly fix it and then go and click, click through and figure out the UI again.
So one question about this, and this is like a literally naive question.
This is not like me trying to be a smart interviewer here.
When you say, like, come up with the specifications ahead of time, like, this is what I want, you know, that test, that API to return.
Are you sitting down, like, in Visual Studio and, you know, and writing out your JSON response that you want to see?
Or are you working with it to develop those specs in the first place?
Yeah, I'm working with it.
I'm just telling, like, hey, I want to send these parameters.
It's all in the prompt, usually.
Right.
Same thing with you.
I think when I hear people say, like, oh, you just need to give a good exit criteria.
I'm like, yeah, but I don't know what that is.
So I just like keep going with my square wheels.
But to your point, you're sitting down with Claude to come up with what the exit criteria would be for it.
Yeah, so that kind of covers the local testing that you do as a developer.
I mean, it's also, as I said, it was so much cheaper to have Claude just go and do extensive testing because you don't have to constantly babysit it.
Can you just go off and try every single scenario, right?
Like manual testing and you can give access to your local Postgres database.
They can go and check records there.
you know, so that helps a lot.
Yeah, 100%.
The second piece is just writing all our E2E tests and integration tests.
It's just become so much easier now.
Yeah.
And I think integrating- Keeping those things stable and trying to keep them not flaky, 100%.
Yeah, it's still a problem, to be honest.
Still a problem.
Like, I've been writing, like, even back in the Selenium Dines days, Cypress, we're using PlayWrite right now.
But it's still a challenge that sometimes when it runs a GitHub action, it's just running on a slower machine and there's a timing issue and you have a complex UX.
So that's still an issue, but we need those things that run on every PR to make sure our system works.
But we can't obviously have like E2E tests for playwrights for every single thing in the app.
So then we have a nightly job that runs and covers the other things.
So it doesn't hold up every PR.
But now we also are experimenting with having these like.
QA bots that just run on the laptop like all day long and they're not using like written playwrights so they're slower so they're more kind of discovering things in real time so they're real slow but they're very reliable because they can like kind of self-heal so if you move a button it can figure out oh the create button is not there anymore like whatever you did a test ID you gave it but it can kind of figure out in real time oh it's not create it's new now Yeah, you just change the neighborhood or something.
Right.
So it can, it runs really slow, but it's a lot smarter and more robust because you can kind of self heal.
A lot of these QA AI tools are out there.
They all claim, oh, we can self heal.
Yeah, they can, but they're all there.
It's more marketing than actual reality, to be honest.
They're still very cumbersome to use.
The ones I've looked at.
Okay, so I get like MCP in a terminal, 100%.
I think we've all done that one.
I get, you know, MCP with Playwright running in the daily, you know, to double check everything and running in the GitHub action for every build.
I totally get all of that.
This last bit, like when you say, like running on a laptop, do you mean a developer's machine or some other way?
We're setting up like a separate laptop.
We're just experimenting with this.
So it's just like checking out the latest build and spinning it up and just like error guessing?
It's more like exploratory testing.
It's like a compliment, right, to our EDE where this bot can just go through our app and kind of look at the story and kind of like figure out if there's anything missing.
So it's not like...
Okay, it's more...
So it's still story by story?
Yeah, but it's more...
Like we don't tell it exactly what to look for, right?
It just looks at the door and it tries to figure out what to test based on that.
Super slow.
We're kind of just looking into it now to see if we can make that work.
Because this is a really nice complement to what we have already.
Yeah.
Yeah, yeah, yeah.
Oh, that's awesome.
So I don't know if it's going to work, but we've done some experiments.
It seems to...
You're going to end up with a rack full of Mac minis doing this or something.
Yeah.
What about debugging?
Oh, when was the last time I opened that debugger?
I can't even remember.
Right.
That's kind of what I mean.
That's really good at it.
Yeah, it's incredible.
Like, especially things that were really hard to troubleshoot back in the day.
Like, you have some race condition, you know?
Like, you could spend hours just pulling your beard thinking, like, what's going on here, right?
Like, you just give Claude as much information as you can.
Any logs, any information that happened at this time with these two users were using the system.
This was the data that was in the database.
Just throw as much information at it as you can.
Most static code analysis, you can figure out, oh, I found this race condition right here.
And bugs, same thing.
It's incredible if you just give it all the logs, all the data, like it can usually just figure it out for you.
I feel like it's faster at reading than I am.
And so that's like, usually if I've got like a production issue, I'll queue up Claude and say, go start to look at this.
This is what I'm seeing.
And then I'll try to like race it and I never win.
But sometimes I will catch the occasional detail that I'm like, oh, I should tell it this.
It might not have seen it yet, but it will still win.
But I still work with it, you know, but it's like I'm a little sidekick.
Yeah, it wins most of the time.
I was so happy the other day or it was a while ago when I did actually beat Claude to the putty.
So I was like, I found a bug while he was looking.
Missed it.
Nicely done.
Okay.
What about integration?
And here I'm thinking about not like your CI, CD pipeline, you know, and that kind of integration, but more like, okay, there's a new partner API or a new product that you're getting access to an API or a department you don't normally talk to, like that kind of integration.
What are you seeing the changes there?
I mean, since we're all one startup, I don't have too much experience with that.
Related to that, like...
When we deploy a new version of the API and then the new version of the UI, we always have to make sure that the UI works with the old version and the new version, right?
Because people are not going to instantly refresh their browser to get the new version.
So that would be kind of similar where you need to make sure what you would use contract testing for and things like that in the past.
There's a lot of things you can do to just tell the AI just like, okay, this is the version running in prod.
This is the version running in staging.
Just basically go look through the commits and just figure out, hey, is this new version going to break the UI that's currently running?
It's found incredibly helpful.
It's caught so many issues.
Yeah, yeah, yeah.
So there's like dependency management there.
Yeah.
Yeah, like just go through, like all our services has like a slash build endpoint.
And I'll tell you this is git commit that's running currently.
So AI will go and fetch that.
And then it can just go into your git repo and figure out, okay, these are the commits that are new in the new build, right?
Versus what's currently running, what new endpoints have been added or any of those going to break, right?
Which is a pretty tedious process to do because you need a lot of...
tests around that.
And if you just have like one UI, you're not going to have that extensive UI test that you test every single functionality for every version and making sure nothing breaks.
So it's caught so many issues where we were like, okay, no, someone just renamed something here or had to do some change here and it's going to break the UI.
So, which is fine.
Maybe that was what we wanted.
We don't want to, we wanted to kind of clean that up, but at least we have to do the deploy off of ours because it might be split second where things are not working properly.
Awesome.
You guys are going really deep.
I'm loving your answers.
Okay.
What about refactoring?
What are you seeing there?
I mean, a lot of people saying refactoring is dead.
You can just regenerate.
I would say refactoring is incredibly cheap now.
I think that's my only opinion on that.
I'm not in the camp that you should just regenerate from spec all the time.
I don't think we're there yet.
I think that's like a maybe we get there kind of thing.
That seems very aspirational to me.
How good a spec would you have to be writing?
Yeah, exactly.
Right.
Like every time, even though you have a really good spec, you're still making decisions telling it what, because these specs is usually very business focused.
And then you have, you make design decision and coding decisions with calling back and forth, telling it, okay, this is how I want to structure it.
I want to use this kind of pattern, right?
That usually doesn't go back to the spec.
So when you regenerate the next time, you got to answer those same questions, right?
Maybe one day.
Yeah, but to your point, you find it cheap now doing refactoring.
So you're doing a lot of it, I guess.
Yeah, I mean, especially, yeah, we missed tech debt before, right?
When there's things we missed or things over time.
Not sort of academic refactoring that I'm hearing.
I'm not hearing like, oh, I think we can improve this kind of stuff.
I'm hearing more like, yeah, there's a tech debt over there that slowed us down.
Kill it.
Make it never slow us down again.
Very sort of pragmatic tech debt approach is what I'm hearing.
A lot of things are on performance, I find.
Like, yeah, you start one way and it kind of.
It gets bigger and bigger and then, oh, wow, we got it.
Like we do how we did this just to make it perform.
And it's just very easy to do that and still have it like, okay, does it produce the identical result to what you had before?
Right.
It's so easy to have it do all those tests afterwards to verify that your refactor is good.
Yeah.
And like with debugging, I find that it's good at finding what the performance problem is really quickly and quickly putting together quick shim tests to see if it actually helped.
Yes, we could have done all of this before, but like it just, it's faster.
Yeah.
Yeah, like I was mentioning before with the whole SQL performance that I got into, that's always been a tool available.
I used to always read like explain plans and stuff like that for really critical things.
right?
Because it takes a long time.
Like you need to go get these sequels and get ready for the query plan and then figure out what all these things mean.
Now you can do it so easily.
And also to explain the query plan, right?
That takes a bit of talent too.
Let's go into that one.
That was one of the articles you wrote.
And if I remember right, you were talking about having built like a skill or an MCP something.
Yeah, it's just a cloud custom command.
You basically give it an endpoint name and it knows how to test that because every endpoint has an integration test.
And it's not so much about like, okay, run this with millions and millions of records in the database.
It's more just about like, okay, do static code analysis.
So if you have your entity framework code, you can very easily go and say, okay, you're joining all these tables.
And then you can even kind of give it, okay, this is how many rows are in these tables usually.
So what the command does, it runs this endpoint in your branch and it records the static analysis on.
on how you directed the database and it checks the query plan and then it checks out your main branch and then it kind of does the same thing and then it kind of compares the two and it says you are overfetching by this many thousands of records you're joining all these tables that are not necessary um and it just gives you a really good side-by-side comparison in like a minute which would have been so tedious before right yeah you basically would have waited for performance to get bad to run like one of those tools yeah you would have loaded up the database with tons of records and you just test it like oh it feels quicker done yeah and that's that's another one that i feel like i i really ought to try out uh god we're going through the list really fast how do you feel like it's changed pair programming mob coding these kinds of things oh it's going to be controversial i think pair programming as we're used to it where you where one person is driving writing commands the other one's looking at the shoulder making sure it's correct and things like that and learning from it that kind of way kind of broke when you started using multi-agents right however very important to have senior engineers sit with more junior engineers to show how you're working with AI right you just assume like oh everyone knows all these tricks I got off my sleeves and works with someone whatnot I think that's important and also like writing the AI workflows together I think is also very important because they're so critical like everything you do now is around a couple of different cloud commands that we always run as part of our process.
So I think that that would be very important.
But the traditional pair programming paradigm where you just sit and watch over each other's shoulder, I think that one needs to be replaced with a different flavor of it at least.
Yeah.
I don't think that sitting side by side for eight hours a day is that productive anymore with Gentica AI.
Right.
So listeners will, may have heard me talk about how the ThoughtWorks approach to pair programming.
It was literally that.
It was like 80, 90%.
You know, you sit down for three hours in the morning and pair program together and you go have lunch and you come back and pair program for another three or four hours.
That's what I did.
ThoughtWorks.
And to be honest, I thought it was exhausting.
As an introvert, I like putting on my headphones and like get into flow and code away and to sit and.
talk to someone for eight hours a day, I was always exhausted.
I have all the value though.
Like it really brings a lot of value.
Yeah.
Sharing the context, right?
Those two people and constantly have a sounding board.
Yeah.
But that's the thing.
Now AI is your sounding board, right?
Like I, right.
I asked AI so many questions every day just to learn and challenge me.
Like, this is what I think.
Like come up with a good counter argument to how I see the world.
Right?
Yeah.
Yeah.
Oh, I love that.
I love that.
Wait a minute.
What are you asking it for counter arguments about?
I mean, any time where I just like, hey, I built this feature gazillion times in the past.
This is how I would do it.
Right.
Challenge me on this.
Is there a better way now?
Right.
Because I've been doing it the same way for the last decade.
Doesn't mean that I guess sometimes you're working on an impersonation feature the other day.
Right.
I built that one.
This is the third time I built that one in a SaaS app.
Like it's like I'm always on autopilot.
I know exactly what to do.
challenge its face.
So I always want to like make sure that Claude is keeping me honest there.
Like, okay, this is how I always done it, but is there better right now?
Okay.
So impersonation here, I think you mean like software as a service where it's like, okay, I'm customer service.
Yeah, like the admin, I want to go in as well.
I mean, it's you and see how the individual user works.
Yeah.
Yeah.
And that way I can service your account or maybe I'm trying to figure out a bug or whatever.
Okay.
Right.
So that's the kind of thing where it's not like, oh, you know, it's such and such method in the base class library that does that.
You're looking for an entire approach.
Yeah.
So you guys have been doing this for a while, you said.
Like what size of team are we talking about?
So six developers.
Yeah.
I started first kind of like evaluate this, how feasible is this to just go all in on this, right?
And I quickly realized that this is really working.
Like, especially if you use that time you gain to reinvest and put these guardrails in and make sure the quality is there.
And the rest of the team has been really great.
The whole team has just picked it up and everyone is just using Claude and multiple agents.
And we have all these shared AI workflows and context files and everything.
So it's...
I think it made it easier that the rest of the team was sold on it pretty quickly.
We didn't have a much- They're open to picking it up.
Yeah.
Yeah.
Got it.
it sounds like you guys are working on developing these things together, right?
Which which makes a lot of sense.
Like I mentioned superpowers earlier, but when I added it to my cloud setup, it started doing weird stuff.
It's like it started suggesting pull requests and I'm like, there's just one of me.
I've never done a pull request.
We're not doing that.
So where did that even come from?
So I can imagine that like a team working together, you really want to do like a ways of working session almost around like what do we want it to do?
So how does that work for you guys?
Are you is that like a weekly tech huddle where you're kind of going through a Cloud MD together?
Like, how does it work?
We have a weekly tech meeting.
We could probably be better.
A lot of this is coming from me.
I'm pushing a lot of these things, but other developers are adopting it and using it.
Right.
I'm trying to get other team members, too, to contribute more to these files, especially people who are really good at certain things.
Like, we got a UX or UI engineer on our team.
They're really good React developer.
I really like to tell it, please.
go and add things in the PR review command, go and add things in the Cloud MD.
Like we need your expertise in here because you're, you're really strong on this.
Got it.
So you're sort of the toolsmith on some level.
You're the one, the agent herder in a way.
Yeah, I'm trying to, but that's, that's where I think, yeah, like a weekly mob session or something where we just like review all these tools on a continuous basis.
That's probably a good idea.
Yeah, yeah, yeah.
We've just been running so fast here.
Like we sometimes we just got to take a step back and like.
How does this stuff even work together?
Yeah.
Okay.
So it definitely sounds like you're bought in.
There's no hesitation in the way you're talking about it.
What kind of increase, what kind of improvement do you think you're getting out of all of this?
It's a lot.
I mean, we started this project the first four or five months.
We were just co-pilot, a lot of manual coding, right?
So we have the velocity back then compared to now.
I don't have the numbers in front of me, but I'm saying it's like probably three to four times the number of points iteration easily.
Wow.
Wow.
Okay.
And using like a similar points currency than that you do today.
Yeah.
Unfortunately, it's changed.
We switched from target process to linear and linear doesn't have 0.5 points.
But we had to like switch everything up a little bit.
So that affected our scoring.
But still, it's a huge improvement.
Yeah.
And that lines up with what I'm hearing from other people as well.
And it definitely feels like what I'm seeing personally, too.
What do you see in the student ratios?
Like six developers, like what's the complement of testers, product people, business analysts?
Like, do you have much of that UI designers or is it all depth?
Yeah, we have a UX designer, which is really helpful.
She does all the Figma designs, which then clients can pull in when it builds a story.
Super helpful.
We have a product owner.
So we have to make sure we write a lot of tests ourselves.
I think it would be helpful to have like a QA that's outside of the team and can use AI, whatever tools they want to just separately from the team, write all these tests and figure out whatever ways AI can be used.
But it's been working pretty well so far, even with just.
the tremendous amount of features we're pumping out every week.
So, okay, so that ratio then, six devs, one product owner, is working out.
Like, do you feel like you could add two more devs without adding another product owner?
Or like, no, you add a product owner first?
Like, how do you feel about, like, how that part works out?
That might be hard.
I think we feel like we're a very big team at six developers with everyone using Identic AI.
Yeah.
Got it.
So to your point, you wouldn't necessarily want to get it any bigger.
Yeah.
It's a lot of value of having a small team, to be honest.
It's just the cohesion and adding more developers is not going to make us go faster.
It could actually make us go slower because then there's just so much like keeping everyone on the same page gets harder.
The bigger the team you get.
If you got a small, tight-knit team, it's very powerful.
Yeah.
100%.
And it kind of allows you to stay small.
Like it takes away any pressure to get bigger.
You can go so much faster.
We're a startup.
We're placing the original product and then we're also building all these new features.
So we're in this phase now where we just like we're building so many new features.
We're going to get to a phase where it's going to be more maintenance, right?
Right.
And we're six developers and we're reaching all our goals in terms of like deadlines and timelines easily.
Yeah.
So yeah, six developers.
It's a good team size, I think.
Yeah.
Yeah.
Yeah.
I'm with you.
Okay, cool.
Has this changed how you hire or staff, how you train or mentor?
We haven't really hired anyone since we kind of went all in with AI.
The hires happened before that.
Like how we mentor people is interesting because it's so easy for anyone now to learn any part of the system just by having a conversation with AI, right?
And we try to keep a lot of these context files with.
very domain specific things in our code base, not just like you ask about code, we can also ask about how the product worked and all these more complicated areas of the application work.
We can sit there and like have a conversation about it.
It's kind of reduced the need for constantly be asking other developers how the system works.
So another question I ask, how do you enable teams?
I mean, that sounds like you're part of your role at least is to sort of figure out like how to use these tools.
You just go to somebody's desk and say like, hey, this is how we're going to do it now.
I mean, looks like you guys are remote because you're in the same remote office you were always in.
So how does it actually work when it comes to it's time to bring the team up to speed with a new approach?
We're actually in an office, but people are only in Mondays and Wednesdays.
And it's a long commute for me, so I'm here the other days.
A lot of it is, I mean, coming up with these shared AI workflows that we can all agree upon and then just getting the team members to use it.
That's been the biggest part.
It hasn't been that big of a challenge, to be honest.
It's just a couple of cloud commands that we're just like using religiously.
And yeah, just having all these like...
So it sounds like you just sort of like push the command into the codemace and then like what?
Show up on Slack and be like, hey, y'all, there's a new command?
Pretty much.
Yeah.
And not forcing it.
There's a couple of things like the PR views.
That's like everyone has to run that.
But there's tons of other commands too.
Like you can use it if you want.
If you want to come up with your own commands, just go ahead.
Like I don't want to force it on anyone either.
Yeah.
Like if you don't want to want multiple agents at the same time in different work trees, I'm not going to force you to do that.
But if you want to, there's a start PR command and it will do all these things for you.
Yeah.
Set it up for you.
Copy all the files.
Right.
But yeah, it's optional.
When it came time to.
choose, I mean, and you've already talked about how Claude is really working for you guys.
Was that a conversation?
Like, did everybody else want to go try out Codex or try out Gemini CLI or Windsurf or something?
Or how did that conversation go?
There wasn't a lot of pushback.
I mean, especially at the time, Claude was way out of Codex, I think.
Now they're more equal.
And people just saw me working with Claude.
what I was able to accomplish.
So it was a pretty easy sell.
So the entire team now has a CloudMax subscription, which I think is incredibly cheap.
I don't know how they can sell it for like 140 bucks a month.
I run multi-agents all the time and I barely never hit any limits.
It's probably just a matter of time before they're going to increase the fees for these tools because we've got to get too addicted to them and then they've got to charge $1,000 a month and we're like, okay, this is still a worker.
Just take my money.
My hope is that Jensen can make the GPUs cheaper and so the Cloud fees can...
Yeah.
I mean, if we can reduce their cost structure, then their profit margins are already going to be huge.
It's just a matter of making sure that the costs aren't too bad.
So we've sort of covered off all the questions that I ask people around like how Gen.ai is changing their work.
And we're kind of getting into the more questions that could be about anything.
Could be about work, could be about personal life, could be about just anything you want to cover.
What is something you've had to change your mind on?
Something that wasn't intuitive at first.
You thought it was going to be this way and it turned out, nope, it was something else.
I think, I mean, obviously just having AI generate the code, I thought that would never like, no, I can't generate all, especially using the cursor early on.
I was really turned off by that experience.
It would go and generate all these files.
Poor quote.
I didn't like it at all.
I felt like I had no control.
So for AI to just write all my code, I thought that's never going to happen.
I'm never going to allow it.
And it's hard for anyone, right?
You've spent all your life figuring out how to be a good engineer and a coder, and then you have this tool that just does everything you spend.
decades and decades learning.
It's, it's, it's, I'm not gonna lie, right?
It's a very hard thing to accept, but it doesn't help being in denial about it.
Um, and this is the way things are now.
So what's something that you're most proud of?
I usually recall this story because it's kind of interesting.
Um, how I got my first job was when I was in university and it was right during the big IT bubble burst.
It was in 2001, I believe.
It was my second year and I somehow managed to land a job when there was like people with 10 plus years experience looking for a job.
And I'm still puzzled how I managed to do that.
And that then turned into like a 20 year engagement with this company where I essentially took over the entire software stack and I was the only developer and I ran this SaaS business as a side project for 20 plus years.
I thought that was, that's a pretty interesting story because I'd still puzzled.
I think I had very high confidence back then.
Because I remember the hiring managers, so I have guys here that with 10 plus years experience, why would I hire some university students, second year of university?
I said something to him, just give me your hardest problem.
I'll just go and solve it for you.
And I ended up doing that.
And I showed him something.
They were so impressed that, yeah, we'll hire you.
And I had that engagement for 20 years.
It was pretty interesting.
20 years, wow.
Yeah, I just gave it up a couple of years ago.
It was...
fairly large SaaS app.
It worked out really well because later on, I was, this was back in Sweden where I went to university and I then moved to Canada and I couldn't work for the longest time because it took so long to get my work permit.
It took like seven months where I was just sitting in limbo, running out of money.
And they approached me and asked me like, hey, could you like rewrite our whole app?
It was desktop back then.
So they wanted like a SaaS app and like, sure, I'll do it.
And it just worked out so perfectly.
I could work for them while I was waiting for my work permit.
I would have run out of money.
And then I maintained this app for 20 years, which was a very interesting experience because very few users and still work on other full-time projects, right, during the day.
Just to see a code base evolve over such a long time.
And I had to do a complete rewrite of a UI, I think, two times.
Just because, like, every seven years, I think you've got to do a UI rewrite.
That's always something I keep in my head.
But the backend was kind of the same.
It was a Java backend.
It just kind of...
got tweaked over time, but it was essentially the same code base as 20 years ago.
Okay.
Hopefully you moved off of something else at some point.
Just guessing.
Yeah.
So it's small updates on the backend, but like big rewrites on the UI because UI changes so quickly.
I think you have to do like almost a quick, a complete rewrite.
Cool.
Every seven years.
That's what I have.
I love that.
So every seven years, you just got to rewrite the UI.
That's, that's good advice.
Yeah.
I'm sticking with that.
Okay.
So what's something that's giving you joy right now?
Outside of work or?
Usually, but not always.
I often interview workaholics.
So, you know, it is what it is.
Yeah, I've kind of fit in that camp a little bit too.
I need to work a little bit less.
One thing that's really bringing me joy, I know you've seen these pictures behind me.
I put these up a long time ago because I wanted to go up one of the mountains here called Mount Temple.
It's one of the tallest mountains in Canada.
So I've been training for that.
And I finally did it this summer.
Took me two attempts.
I failed the first one.
because it was too much snow, even though it's middle of summer, we just got hammered with snow and it got way too dangerous once we were like 70% up the mountain.
But then I came back a couple of years, or sorry, a couple of weeks later, and it was a beautiful day.
It was no snow, just a little bit on the top.
And yeah, it was a 10-hour hike to get up and down, but it was incredible.
And I just come back from Utah, where I was last week hiking in Zion National Park, which was an incredible experience.
The outdoors bring me incredible joy.
Cool.
I have not been to Zion yet, but Monument Valley in Southern Utah, just beautiful down there.
Wow.
Holy cow.
Stunning.
Yeah.
Awesome.
Yarnas, thank you so much.
This has been great.
Thank you for having me on.
They're small, highly maneuverable, and in some respects, they think for themselves.
Robots.
Exactly.
Cool.
I'm going to pause for a second.
Can you hear my dog barking in the background?
I used to record in a studio, so now I'm doing this out of my home studio.
What am I going to do?
I'm supporting.
They're both rescues, but it's always the one that's barking.
She's kind of got a jack-rissons-herior kind of thing to her.
She stopped.
I'm going.
I think my wife just got home.
I think that's...
part of why this is going to...
I'm gonna try to just keep going.
