# Agentic Engineering: Scaling AI Coding in Enterprise

**Podcast:** Engineering with AI
**Published:** 2026-08-17

## Transcript

The conclusion I've reached, though, is that if we looked like we were accidentally creating greater quality, we would respond not by achieving that quality, but by lowering our standards or increasing our haste until we managed to lower quality back to where it was before.
Because the level of quality we have in software is not the maximum we know how to achieve, or at least I don't think it's been in any codebase I've been part of.
It basically lowers.
The conditions will accept.
And if it gets lower and things start breaking, then people will get it.
So it's just above the waterline.
And in engineering with AI First, we're having a guest back for a second time.
Most of our audience will remember Chris.
When he was last on the show, he led with Human in the Loop is Dead.
He's also been one of the most often downloaded episodes of all time.
And I don't think it's just because he said something spicy.
I tried that with other guests.
There's something else to it.
We wanted to have him back today because of the book he's been writing.
So let's get started.
Chris, what's it called?
What are you doing?
What are you saying in this book?
All right, so listen to me back.
Okay, so the book that I'm writing is called Agentic Engineering at Scale.
And basically the concept is I wanted to talk about all the stuff that is beyond personal productivity.
Like there's a lot of personal hacking going on, but then a lot of people want to know how to get stuff done in groups of people or communities of people.
Maybe I can kind of say how I've...
divided things up in the book because I want to make it work for people who have real problems in real organizations, not because I don't find it really fascinating how you set up your right agents MD or exactly how you code as an individual.
But, well, selfishly, that stuff's changing really quickly.
So very hard to put into a book in a way that will not be obsolete in a month's time.
But secondly, I think the gap.
that a lot of people I'm seeing ask are people who come in after hacking on the weekend, getting all sorts of stuff done, and then find that they just don't know how to make that thing or that magic happen or apply in their workplace.
So that's what I'm trying to tackle.
Okay.
So this is not so much for the person who's bi-coded something and now wants to get it in the app store.
This is more for the plucky engineer who wants to help out the team he's on, he or she?
Yeah, so I guess the, so as I'm a consultant, I work for ThoughtWorks.
And so often we give advice to our clients about how to do better in their software engineering.
And that normally means because they have to be big enough companies to want to pay externals to give them opinions, that they try to get large things done with groups of people coordinating.
And so that doesn't, it doesn't decompose into just everyone go off into their own corner and vibe code their own app.
There has to be.
They have to care about quality that maybe people don't care about as much when they're prototyping.
They have to care about managing priorities and specs and things like that.
They have to worry about teams of people working together.
And maybe the biggest thing, which I opened the book with, which I think is a bit of a contrast, is they have to worry about legacy code.
So a lot of people you see on LinkedIn or wherever, they're basically talking about things that they started last night and maybe you've got 100 GitHub styles already.
But they're not thinking about, here's the system that pays.
a thousand people's wages, how do I work with that in a way that respects it and doesn't break it, basically?
Yeah, yeah, yeah.
Okay, cool.
What was your process like for the book?
Did you research and talk to people and let's hear that part?
Well, selfishly, the fortunate thing is that my job consists of coming up with opinions in this area and working with clients to solve their problems.
So I do get a bit of synergy out of my personal thinking time and my work.
So a lot of what I write about is maybe a synthesis of my own experience across many different clients, which is one of the privileges of being a consultant that you get to dip your toe into many different waters, but also the collective wisdom of my colleagues at ThoughtWorks.
So I do a lot of sharing of drafts and asking questions on the internal chats and things like that.
I mean, of course, also I do a lot of reading to inform or try to make my book semi-informed as well.
But really it's trying to channel the...
experience that people have already started to have in this area.
Yeah.
It's no easy thing writing a book.
There turns out to be a fair amount of work and writing down, publishing down, all the stuff.
Yeah.
Well, and especially, I guess, fascinating now is I don't know whether people will be starting to assume that books have been written themselves by AI.
I've certainly had at least one article that I wrote by hand a while back have someone report to.
Reddit, I think, for being an AI.
I'm not quite sure what caused that, and I don't particularly blame someone for thinking that because I guess there is a lot of mixture of AI content or AI-inflected content or something out there, but it is kind of an interesting thought that I might be just putting quite a bit of effort and quite a bit of part of my life into compiling 100,000 words and then have someone assume that they...
Yeah.
They could do it themselves with Fable.
Yeah.
What was it?
A two, three sentence prompt?
Basically.
Yeah.
That's all it was.
Yeah.
No, it doesn't work that way for sure.
It's funny though, to your, to your point, there are certain tropes that like just hit me in a certain way right now.
Yeah.
Like when.
If what actually matters here is, I'm like, okay, yeah, that was probably, you know, it was anthropic model.
It was Claude that you were talking to that generated that.
And, you know, everybody at first was trained to look for the em dash instead of the hyphen because, like, who even knows how to type that?
But now it's certain other things.
But now I wonder, like, is it inflecting our speech?
It's picking up on what we are doing, and then it's doing it, and then we're going to follow it.
I don't know.
So which came first in all of that?
Yeah, well, I have actually an em dash problem because the way I found the easiest to write is dictation.
Maybe because through conversations like this, I find it easier to maintain the flow while speaking.
Yet while speaking and then reading a transcript, in speech, we have a lot of meaningful pauses that lead on to the next thing we say.
So I've naturally, when I get my raw speech transcribed, it has quite a few em dashes based on the software I'm using.
So I'm now faced with the paradox of, Do I leave them in because that's authentically what I came up with?
Or do I try to write some kind of AI tool to get rid of the end ashes to pretend that it's more natural than it actually is?
And so I'm a bit confused as to how I should do that at the moment.
100%.
I'm certainly doubting myself.
These are the problems that we now have as a society to solve.
Yep.
Cool.
Obviously, you don't want to give away the ending or, you know.
That's not a murder mystery.
It is okay.
I can kind of maybe engage with the key arguments of it.
I think if someone got to the end of a technical book and only kind of found out right at the end what it was, I'm not sure they'd actually invest as much time.
Yeah.
Well, but I'm also, I'm wondering if there's a story in there that, you know, to your point, a lot of this comes from stories, real and synthesized.
Is there a story you would tell out of this?
I think there's, I have a narrative of order of concepts.
It's a big.
overarching theme is that control theory is a really useful way of thinking about how to get coding agents to do good work and that if you're going to take that seriously it's not just a matter of an individual coming up with the kind of the guards guardrails checks sensors and things like that it's actually something that you have to start sharing across your organization which leads to a whole bunch of problems because we already found it kind of difficult to do platform engineering, but I think I have become convinced that if you're going to achieve with agentic engineering at scale, so just getting that book title in there, another kind of branding plug, you will have to start thinking about sharing common expectations of what good code looks like in a way that coding agents can access at scale.
And then you'll have to worry about the knowledge management and the human collaboration that gets to the consensus on those artifacts.
Right.
A team that I serve has taken to having a confluence page that sort of is a little bit of a starting point around like, OK, please don't just go write your own crazy Terraform willy nilly.
We use a YAML file that our platform team knows how to ingest.
Here's the ropes, there's the bathroom kind of stuff.
And that way, since it's in the confluence, the first thing we do in CloudMD is go read it.
And it's like, OK, now none of us have to tell it again, you know, and we can kind of keep it up to date.
So I'm imagining that kind of thing, but you may mean other more nuanced things in addition.
So I divide up the different kinds of, and I'll say harness for now, and maybe you can substitute in whatever word is popular at different times.
You know, we can refresh this as we get different jargon as we go along.
But I divide harnesses into four kinds, and I think you need a balanced diet.
So maybe I think the two most important distinctions of different kinds of harness is whether it's Something that tells you yes or no.
So whether it's giving the information back to the agent and the agent has to make its own decision or whether it actually gives a red or a green signal.
So a unit test tells you, no, it's not good enough or yes, it is good enough.
But something like using play route from the browser gives the agent access to look at the site you're working on, but it doesn't tell the agent whether it's acceptable or not.
The agent has to kind of apply the judgment.
So I think that's kind of, that's one of the distinctions.
The second distinction, just a link to what I said about control theory, is feed forward versus feedback.
So maybe I should unpack that because feed forward is a somewhat unfortunate nilligism.
But feedback is, I've done a bit of the system.
Let me use this to perceive it and work out what's going on.
Either to tell me whether yes, no, if it's a judgment-based thing, or just to let me get information or context about what's going on, kind of observability.
So that's kind of feedback, which people are familiar with.
Feed forward is effectively guidance.
So it's, you know, in control theory, feed forward is the stuff that tells you where you should be aiming for without you having to give it a go and evaluate where you are.
So if you kind of think of those two things as axes, because I'm a consultant, everything becomes a quadrant, that you get kind of things that are normative, tell you yes or no.
that are feedback, things that are normative that are feedforward, and you get the informative versions of feedforward and feedback as well.
And so that combination, I think, is really important.
So to your point about guidance, I would call that a guide if you have something about the golden path in your confluence.
So that's informative feedforward.
So the agent can use that to do a better job.
It's not going to tick off the agent for getting it right or wrong, and it's not going to allow the agent to evaluate what it's done, but it is actually one of the important things.
in order to have a good system of input to the agent.
And I would say a lot of people who are struggling at the moment tend to be over-indexing on one kind of harness or the other.
So they might be trying to lean entirely on feed forward.
I'm not going to give the agent any way to mark its own homework or understand what it's done.
I'm just going to tell it what good is and hope that it listens.
And it might listen, but it might also go off track and not be able to self-correct.
Alternatively, I see people focusing far too much on feedback.
which leads towards brute force kind of situation.
So they don't give the agent guidance on what a good design looks like.
So they don't give the agent a design system or they don't give it the templated infrastructure for your company.
They just say, go for it and see how you go.
And that results in cycles, which is good, but too many cycles and it doesn't converge on a good answer as quick as it would if you'd started near to where it should have been.
Right.
And I don't think these things learn on the job in the same way or as quickly as humans do.
So, yeah, that sounds like it could be particularly painful.
Well, yeah, learning on the job or just online learning in general, some people would say is one of the big architectural problems with the current generation of LLMs.
It's actually, you know, it's not what the everyday public thinks is happening.
The everyday public assumes that if they ask a question about LLM, the LLM will learn at the same time because that's what we do.
But as we know, if we want to...
feed information we've learned back in, we have to do it explicitly with some kind of agent memory mechanism.
Exactly.
Yes, we have this business of memory saving now and those kinds of things, which is kind of trying to approximate it.
I'm not sure it's exactly the same thing, but because that just results in more feed forward.
But maybe that's fine, I guess.
But you've got this concept of like the context window is only so big, though, right?
So I assume all of this feed forward is eating into that budget in the end.
So I guess progressive context disclosure.
Maybe an overly bureaucratic term is what people usually say for trying to stage the input so the agent doesn't just read the whole world through a straw.
So, for example, if you're giving your agent ADRs, you should be giving it an index page with clear kind of titles so that it knows when it should actually go one step further.
And when I've seen people profile the context usage of their agent accidental.
over reading of context is probably the number one thing that burns tokens.
They didn't really realize it, but they've got a kind of a series of hyperlinks that the agent is preloading because it's been encouraged to go from the agent MD to this file to your confluence from other page on confluence to a manual.
And it does that maybe on every session or almost every session.
And that causes a bunch of irrelevant stuff to junk up the context window, like you say.
Yeah, yeah, yeah.
And it gets dumber as the context window fills for sure.
So this is to be discouraged.
I don't know about your experience, but I think the common figure that people say is that if you're more than 50% full, it starts to have trouble with knowing what to pay attention to.
So you start to get erratic results.
Yeah, right.
I often think of it as it doesn't know what to forget to keep room in them.
But yeah, you're right.
It's also just the problem of what should I pay attention to.
This thing's got so much in it now.
Yeah.
In my mind, it manifests in maybe early on in your Agents MD or even in your session, you told it, please remember this, this, this.
Do not do this, this, this.
And once you get more than 50% full, you start playing Russian roulette with whether those instructions are still going to be brought to mind at any one time.
Yeah.
Through the magic of time, the time travel that podcasting brings, people listening to this episode, if they listened to the last one, it was with Birgitta Buckler, also at ThoughtWorks.
Yeah.
And she talked about how we say these are guardrails.
Those are the things that keep your car on the highway.
But sometimes your car just goes straight through and the guardrails are just gone in this world.
So are they guardrails at all?
Yeah.
Feedback and feed forward definitely seem like better words for what's happening there.
And to be clear, I partly got those terms literally from Brigitte, who I talked about this before.
So I don't want to claim to have superseded or improved necessarily on her thinking.
I think she's done a really good job of distinguishing between different kinds of things that we have in harness engineering and spectrum and development.
Definitely.
She's given us some more words.
This is so exciting about your book, Chris.
What can you say about when it comes out?
And is this going to be just Dead Tree?
Or I don't know.
Will there be a movie made?
Like, do I get to play the starring role?
Probably not.
Yeah, I mean, I guess a dead tree is not the only format that you would release a book in these days.
I do have a promise of that being the case.
I have an O'Reilly animal for the front cover, which is Australian fauna, which I asked for.
They didn't agree to get me a deadly Australian animal, so they found one of the cute ones that is not deadly.
But still, that's pretty cool for kind of an eco-nationalist like myself.
It's supposed to come out in Q1 next year.
I'm about to say something that I'll probably regret.
I'm a little bit ahead of schedule in terms of the writing, though I've discovered it's a really good embodiment of the difficulties with improving the speed of coding, not necessarily improving the speed of the machine.
I'm working in a collaborative effort with a bunch of people who do other things in their day job with publishing and they have other schedules.
So me dictating or typing slightly faster doesn't necessarily alter.
when the end product is, much like using a coding agent doesn't necessarily get you live faster in a real project.
But yet the schedule is Q1 next year, but the early release is in the works at the moment.
So that'll be on the O'Reilly platform in electronic form with, I think, the first four out of 12 chapters.
That's amazing.
Cool.
Aren't you going to keep us waiting to figure out which Australian charismatic father?
Well, I just assumed you wouldn't know.
So I really like to use the obscure ones.
The animal is a honey possum, which is basically the size of a mouse, but it has an especially long nose because it's designed.
Well, I was about to say designed.
I didn't mean that.
It's iterated towards being able to.
suck out honey from banksia and protea flowers, which are these kind of like enormous flowers of Australian flora.
So, yeah, it looks pretty cute is the idea.
And, yeah, I don't think I know of any fatalities at all involving them, but I guess...
Involving honey possums.
Yeah.
Yeah.
Tiny possums, basically.
There's an analogy in there somewhere, the long straw, something, I don't know.
Yeah.
Yeah.
I think I've learned, obviously, I've done a lot of technical reviewing in the past, so I've been...
review of a lot of books, this is my first time actually being the driver rather than the backseat driver.
But one of the things I learned is that publishing companies want to be publishing companies, not people who spend all their time explaining to adults why they can't have the elephant or whatever.
So in the end, they don't try to come up with really direct mappings of cover animal to subject matter because that ends up being, just like coupling in software, that ends up being...
Yes.
A bit onerous, even though it might seem neat at the time.
Yeah.
A bad idea for your publishing process to try to find.
Exactly.
If you change what chapter seven is about and they have to get you a new animal.
I don't know.
Right.
That too.
That too.
This is cool.
Super exciting.
So I'm wondering, did you check out the radar?
What were your thoughts?
It was a lot of fun for me to go back and review what you and other guests had said and put this together.
Yeah, what were your thoughts?
To start with, thank you for putting it together because I think my kind of meta observation is that we're using words in really incommensurate different ways.
So your work you're doing of kind of trying to organize the thoughts to actually start the process of us at least being able to even argue at each other rather than past each other, I think is a pretty important function.
So thank you for doing that.
I think for me, I was quite attracted, perhaps because I was interpreting things through my own lens, the things that were, I guess, talking about feedback in some way.
I saw a lot of things through that lens, and that may be to tie it back to the work I was doing of trying to reconcile feedback and feed forward, is that I can't say I scientifically remember exactly every portion of things, but I wondered whether we were focusing too much on the inner loop and not enough on the guidance of...
how to get too close to what the answer is before we iterate.
So I kind of, I liked it both for helping me to explore the different ways people have thought about those things, but also, you know, I was kind of, in a way, using it to map out what parts of the conversation we haven't had yet.
Got it.
Now, when you say inner loop versus guidance, I'm not sure I can exactly, like I could make three guesses as to what you mean, but help me fill it in.
Well, I'm probably not even using my own terminology in a standard way, but what I meant specifically was the experience of an individual developer.
They've got the coding agent, they've got things set up, and they want to give the agent a way to tell whether it's doing a good job.
I think I seem to remember quite a few of those.
So in my mind, that's actually one of the most useful quadrants of the four, but it's feedback and normative.
So how do I get a linter?
to tell the agent that it did wrong and it has to fix it?
Or how do I, you know, how do I otherwise control the agent?
I think that is really important because anything that's in there is part of the discretionary process that you can set the agent off to do.
To tie it back to the earlier comment, I'm not sure I phrased it exactly this way, but a human in the loop being dead.
But if you want the human in the loop to not be there, you need to give...
enough scaffolding that the agent can try, fail, try better, fail better, succeed.
And a lot, I think, of our attention has been naturally at that question of in the small, including your guests, because that's a problem we have to solve before we can even earn the right to encounter.
other problems about how we get to good results.
Yeah.
I'm seeing interesting patterns in terms of how to make that safer, right?
And whether it's like Swamp Club, which I learned about only a few days ago and still need to dive into.
But it's apparently aimed at sort of DevOps scripting and helping the agent become, well, repeatable.
Essentially, I assume that this means it's going to become a tool that the agent uses instead of letting the agent GPU it.
Yeah, yeah, yeah.
way a different time.
But yeah, tools like that, tools like code mods, you know, like building something that goes and does the refactoring for you so that it doesn't have to imagine how to do it each time.
These not only save on tokens, but also make it safer for us not to be in the loop, I think.
Yes, I think if don't use inference to solve something unless you have to, is I think a reasonable rule.
Although I must also say that like, I think some of our enthusiasm for these tools is partially due to wanting to retain the kind of control we have when we use tools like refactoring tools and compilers.
So I've certainly heard a lot of people talk about how they think it's a problem that agents are not deterministic.
And so their solution to that is, well, how do I introduce determiners as much in the process?
So I completely agree, actually, that if a codemod can do it, that's a more efficient way of doing it and it's a less error-prone way of doing things.
But I don't, I think...
It's a strength in a way that agents are non-deterministic.
It's just a strength that we are not used to exploiting.
Right.
And when I'm working with organizations and I'm trying to make sure that the engineering staff in that organization has access to good models, I often find myself saying, I just want my smart friends to have the good tools.
Yeah.
But I feel like this also applies to the agent.
I want my smart agent friends to have access to the good tools, whether it's Swamp Club or CodeMods or...
linters, obviously, but sort of all of this category.
And I guess these are all in the feed, well, many of these, not all, are in the feedback side.
But some of this, is it almost something different?
Like, if we're expecting the agent to reach for that tool in the same way that it would reach for LS or Bash or whatever, is that feedback or feedforward?
Or is that something in the middle?
So I think it might be, I don't want to introduce three dimensions because...
Even in electronic books, you can't really represent it.
But I think the kind of mechanistic versus the inferential is a little bit of another distinction because I think you could introduce it for other things in terms of the way the agent invokes it.
But I think that a lot of the advantage to the – it's not just that it's mechanical.
It's that you'd need the agent to do less work if you can express the possibilities of what you want the agent to do and what effectively is a – smaller language or a smaller set of levers.
Basically, you're asking for less information to be outputted by the agent to make a good choice.
So you're not asking it to wade through a billion invalid possibilities to find the one good one.
It kind of matches up with the idea of the golden path from Netflix and that you mentioned earlier, which is that if you give it a more constrained set of options of how it can act, you're putting less of a burden on it.
You're making it more likely it'll choose well.
I think I'm learning that the story I told about GoToConfluence to find out how we deploy is something Netflix talks about called the Golden Path.
Do you want to tell that story?
I think it was originally they taught the concept of the Golden Path.
So Netflix were earlier adopter of microservices, expanding very rapidly, aiming to have a really high talent density.
But they also then realized they didn't want everyone to be spending their high talent density in.
building infrastructure from scratch.
So how do you square the kind of anarchist freedom principle with the kind of centralization urge to have consistent platform?
The way they did it, and it's obviously been copied by others because they're an admired company, have done many interesting things with tech, is that they have a golden path that says, if you're using Java and you're using this AWS service, then we support that really well.
We document it really well.
But it's not you can't leave the path.
It's that if you're on a path, you get first-class service across all of our tooling and other things.
So it's a way of, in a way, lowering the team's cognitive load because they don't have to evaluate every container running service that AWS has for themselves.
They know that this is the one or these are the two that Netflix supports and that they'll get great tooling out of.
But if they really do have a unique need, they're not constrained from...
going off-road, basically.
Yeah, yeah, yeah, yeah, yeah.
I can remember these conversations coming out of Netflix, you know, sort of 5, 10-ish years ago.
Like, I think I was in Cincinnati, so that would make it 10 years ago at a guess.
And yeah, that sort of conversation about, like, yeah, you can use, you know, this framework instead of that framework, but, you know, as you say, it's not the golden path that's not as well-supported.
And that's a neat way to split that.
It seems like, interestingly, in my own observations with my own clients, We're kind of not doing that with AI.
And maybe that's like a little bit of a fear about the non-determinism and security impact.
I am aware of companies that have already seen, well, I mean, we've all had the Shai Halud Worm, you know, orders, you know, come in and inject scripts into our systems and so on.
So I get why people would, you know, let's not go with the anarchy path this time.
But I'm just dreaming this up right now.
Does that match your experience?
Or do you see organizations which are trying to, hey, let's experiment, let's allow experimentation, but this is the golden path versus what I'm seeing, which is more like, no, this is the path, you know, blinders on.
So I don't think organizations have got to grips how they can, as a collective, exercise any kind of opinion one way or the other, is my experience.
So teams.
build their own skills and share them among the team is actually the most common level of collaboration I've seen with exceptions of maybe some specialist services like token gateways or whatever.
Most of the time, there is no collaboration of agents working to do things.
So it's the equivalent of completely atomized teams.
The most I've seen is some companies do have systems for sharing skills.
So someone will come up with a really good code review skill.
And then they'll share it with other teams.
Most commonly, it'll get passed, maybe unofficially.
There are some organizations where there's like a central repository for them.
But even at that stage, and we're talking about really basic stuff, not infrastructure templates or design systems like you're kind of hinting at with the golden path.
They're just, here's a way to review a pull request or whatever.
They get put on a central place, but it tends to rot at about the same rate as a wiki page as well.
unfortunately.
So in terms of the choice, does the organization say agents must code in exactly this way versus encouraging the agent?
I think actually by default, we're not restricting the agent at all, not because organizations have decided that they wish to see self-organization reign, but just because the governance functions of technical organizations haven't got traction on how to exercise their function with agents.
Now, I think actually we're going to see that happen and in a very clumsy way soon where, you know, there will be a set of guidance that your agent is expected to load and the CTO will be able to like change it on the weekend.
And where previously the CTO might have wanted you to start writing unit tests or move only to outside in tests or adopt this latest thing that they've heard, they previously had to kind of rely on this chain of influence of telling that.
head of architecture and doing all hands.
And then if necessary, people could just ignore them and get on with their life.
But I think technical authoritarianism will actually be a risk of having shared harness artifacts because it will be possible to directly inject, you know, the prompt injector will be your CTO, maybe is another way of putting it.
And if they have great judgment, maybe that's good.
But I can't say that, you know, being senior always gives you infallible tech judgment necessarily.
The immediate flashback for me is working in one of those organizations where the entire company, and in this case, 300,000 people, the exact same Jira template across all projects, whether they were accounting or finance or software development.
And yeah, so authoritarianism gone wrong.
But maybe this will be better.
Maybe.
Fingers crossed.
I think some level of standardization is good.
That is actually the Netflix inside, right?
It's better that kind of we're a Java shop or kind of we're a .NET shop or kind of we're a Ruby shop because that does help us collaborate more effectively.
But the excessive standardization is kind of pretending the world is simpler than it is.
I think it is very relevant because we're seeing ways of scaling decisions with agents that we haven't seen before.
And so, yeah.
stomping over the local resilience of people taking context into account and localizing decision with that, that might actually be stomped on a little bit by having organizational level harness engineering.
And so I think we will see over correction and then re correction in that area.
Cool.
So, yeah, I, one of the things you, one of the things that you responded when I asked you about the radar.
You said, well, Kyle, we didn't hear your voice.
And true.
That was somewhat by design.
It's funny, though.
On review, when I was doing the work of putting together the radar, I thought to myself, because what I did is I fed the transcripts into Claude.
And Claude comes out and says, you know, Ahmed said this and Ahmed said that.
And Chris Ford said this and Chris said that.
But it was giving them to me by episode.
But I'm not sure it actually knew who was speaking.
So some of the things that it would attribute to a guest, it might have been me that said it.
And so like certain things that I recognize is like, yeah, that's me that says that.
It's like AI is an amplifier.
I'm not trying to claim that one.
Like probably somebody said that to me first.
But I know you are the bias in the model is kind of what you're saying here.
Yeah, no doubt.
No doubt.
So there is definitely some of Kyle in there.
And I think that is one of the things.
If your team is fast at making messes today, it's going to be faster at making messes tomorrow.
That's definitely something that I said more than once.
Yeah, thanks for checking out the radar.
Thanks for being part of it.
Appreciate that.
I'm trying to think about what's changed since you were on.
It looks like it was back in April.
Yeah.
And I guess, you know, you were saying when we prepped the March of the models.
Right.
Fable and the paid world.
And of course, like Kimmy3 and others in the non-paid world.
Yeah.
I mean, it's been it's been an interesting time since April, for sure.
Yeah, I would say that the number one thing that's changed for me in my professional context is that the amount of people freaking out over costs has ramped up.
So, you know, we did have that faddish.
idea of token maxing at one point, right?
Where maybe literally consuming more of a resource was considered virtuous, so we'll have a leaderboard.
I don't think that lasted very long, and I guess it was a little bit of a stunt.
But we're seeing, at least I'm certainly seeing in my professional life, organizations go the other way.
So, you know, unlike us, when we have our personal subscription, organizations are on paper, token, API plans of different kinds.
And even from the extreme of How do I make sure that no one in my organization accidentally spins up a recursive Sorcerer's Apprentice agent job and costs me $50,000 through to just widespread reported people having had a budget set and then blowing that budget?
Roughly, I think quite commonly in one week out of a month is about the common ratio between what people expected and how long it takes.
This is causing a lot of consternation because...
people have remembered that we don't know how to measure software productivity, software development productivity.
So when the costs were low or maybe it was a few teams who were trying things out and certainly they might have been prosecuting their C-suite's agenda because the C-suite was interested in getting better at AI, it was kind of fine.
But as more people blow more budgets, people who were responsible for being able to defend ROI of things.
are finding that they can't easily work out whether it's worth paying that much in tokens or not.
And so that's a little bit of a kind of parties over moment in a lot of places, which is interesting.
I can remember hearing, you know, C-suite folks from various AI companies going on podcasts and saying, you know, if your engineers aren't using so many tokens per day, you know, you need to worry about that engineer's productivity.
It's like, wow, that was brilliant, you know?
I mean, on some level, there is maybe some value to it.
I mean, I kind of get it.
In our own organization, you know, so-and-so has a reputation of using a lot of tokens and no one's sure why.
And other people, it's like, well, you're not using any, so are you learning fast enough?
It's not a useless metric, but certainly, like...
Like you say, we don't know how to be healthy about our metrics yet, I don't think, in software development.
So all metrics can go toxic and problematic.
But it really has been interesting seeing that swing around.
You know, people like my mother is sending me articles saying, oh, AI is getting expensive.
Maybe you should worry about getting into AI now.
It's like, well, yes, but.
I mean, lots of professions, the equipment is more expensive than the wages.
And in those professions, people argue about you'll raise less.
So the problem we have as engineers is we have the biggest cost.
And while we're the biggest cost, it is worthwhile for people to argue the most about whether there's a slight increase in salary.
If we did get to the stage where Jensen, who I think you're kind of alluding to, you know, is having people using hundreds of thousands of tokens a year, then it becomes...
The more we worry about the token cost, the less proportionally people might worry about the human wage cost.
So I'm not sure whether it's a dystopia for those of us who are already skilled software engineers, though it might make the economics weird for others.
Yeah, yeah.
Well, and you raised the Jensen Spectre.
Well, to your point, I raised the Jensen Spectre.
But to name him, I keep thinking to myself, as long as he can make this cheaper, as long as he can keep making the chips better for less money, you know, then we are going to be fine.
Yeah.
Which almost kind of brings me to another question.
So I've had this analogy bouncing around my brain for a little while, and I haven't managed to talk to anybody about it.
And so I'm going to run it past you.
That's good.
Just between you and me and...
Yeah.
Thousands of listeners.
Yeah.
Yeah, yeah, exactly.
And it is the analogy of FedEx and the fax machine.
And so some listeners will know the story, but for those who don't, I'll tell it.
It's my understanding that when the fax machine, the humble fax machine, right, this essentially black and white printer combined with scanner, combined with modem that, you know, can sit on a table and can be plugged into a phone line and plugged into a power line.
they can now call each other and send documents back and forth.
And you can drop, you know, a letter in the scanner side and it will scan it and then it will call somebody up and it will get printed remotely on the other side because print jobs are essentially data.
And so we can move those over phone lines.
And as I understand it, FedEx, when they first saw these, thought, this is great.
You know, Chris is in Barcelona, Kyle's in Toronto.
We could put a fax machine in each office.
If Chris wants us to send something to Kyle really, really quickly, we can fax it.
And it won't look as nice.
It won't be the original.
But we can get it to Kyle for so much less.
And like, we don't have to charge all that less.
And when fax machines first came out, they were 10 grand a pop.
So this maybe seemed like it was going to be a thing.
But what they didn't realize is that they would end up being 80.
And so pretty soon, everybody would have a fax machine, you know.
And big companies almost immediately realized, well, yeah, we should just do this instead of paying FedEx to do this.
And eventually, medium-sized companies, small companies, people.
And then eventually the internet and it was all over.
I wonder if this might apply to AI.
You know, companies like TinyAI are building these tiny little things that you can now run a reasonable model in.
And the better and better the algorithms get around compacting models down and reducing the amount, making it more efficient for tokens to do their work or for models to do their work.
Yeah.
Like, will this be the kind of thing that, you know, we're just going to run it locally?
At that point, that's really going to change the economics of this entire game.
So what do you think?
Like, am I dreaming?
Like, what do you think?
Well, I can certainly come up with a direct parallel from what I've seen.
So the problem there was that FedEx were thinking about competing against their competitors as though their competitors didn't also have fax machines or also their customers didn't also.
So kind of imagining a world where they compare everything else saying static and then moving forward.
I do actually see that quite a lot in the software world.
kind of simultaneously going, oh, code is cheap with coding agents.
That's great.
But they're still pretending that if that were true, that code is valuable.
And that's not an equilibrium that can hold.
So, you know, companies are thinking, oh, okay, I can have higher margins because I can do what I used to do much more cheaply.
That's awesome.
But they're thinking the fallacy that they're operating in a static world and they're comparing their offering to what their competitors or substitutes were able to do before.
coding agents became popular.
So I think there's definitely a lot of like equilibrium settling that will need to happen.
It would definitely be ironic though that if the, you know, if the large LLM providers having, you know, done the R&D and received a lot of capital to eat the world with LLMs were outcompeted by locally artisanally homegrown run.
things on everyone's laptop.
Yeah.
Once the models are good enough and fast enough, once they're smarter than us, you know, like, and, and we can run it on, you know, our phone, then does it matter if it's 10x smarter than us or 100x smarter than us?
Like, maybe.
And this may not be for a while.
I mean, things have changed fast in the last two years, but maybe not this fast.
But, so yeah, if, but if it did happen, if it did get to the point where most people could get just fine results, on their phone or on their PC, on their Mac laptop, whatever it is.
I don't know.
Like, I wonder what we're paying other people for the tokens for.
Yeah.
I guess I will adjust fine remain static because that makes sense a little bit like the thing that Bill Gates didn't say about, was it, I don't know.
640 kilobytes of RAM.
It should be enough for anyone, right?
I don't think he actually said that.
But that's kind of saying, oh, okay, well, that amount of intelligence should be just fine.
I can't see why anyone would need more than that.
And I think I do have a little bit of, faith slash pessimism, I'm not sure what emotion to have, about our ability to find ways to waste, consume, at least desire the additional capacity that exists in the world.
I think the idea of us going, yeah, that's fine, we've had enough, unfortunately, is not something that our civilization has done very often in the past.
Well, certainly not with whatever the hot thing is right now.
Whatever the thing that used to be hot, we want to optimize the heck out of that thing.
I thought you were talking about climate change, but you would know you were talking about more metaphorical, but yeah.
That's a whole other podcast.
Yeah, fair enough.
I won't try to sidetrack there.
But yeah, I think that there will be desire to increase consumption and things becoming cheaper increases consumption is the...
Jevons paradox that has suddenly become hot again, right?
So it'll be interesting to see.
But maybe there's a stage that we could consider in between the two of running locally and using the frontier model providers, which is there are a lot of companies that are using their cloud hosting providers to run models for their own in-house inference as a way of either an insurance policy against...
rising token costs from, you know, Anthropoc open AR Google, or even just a way of trying to do it cheaper because as long as they have the ability to capacity plan and manage utilization, they can just take the margin of the token provider out of that.
So I think companies are definitely making, they're trying to make sure that they have levers so that they aren't dependent on a supplier that might decide to massively hike its costs.
Yeah.
Well, or, foreign government.
Like, we had a guest on, Ahmed Misbah, and he's working on Arabic LLMs, right?
And I think it's a couple of things.
One is, you know, the language and being able to, there's an interesting computer science thing about what happens when we teach at other languages besides English.
But I think more to that, it's also just sort of sovereignty, you know.
And, you know, Canada, where I live, we have one of our telcos is standing up data centers.
you know, with NVIDIA GPUs in them.
We're not creating our own GPUs, but the idea of a sovereign AI, we do have local model companies.
You know, there's a Canadian model.
I'm not going to remember the name.
That's going to pain me.
But yeah, and so the idea that we can't all be beholden.
We saw what happens when the government of the U.S.
decides to tell Anthropic, I'm sorry, your model is too something.
And so you have to turn it off now.
I'm not sure we all want to be beholden to that in the end.
If I understand correctly, there's rumblings in China of actually maybe not wanting people to use the Chinese models as freely in the West, the way that's happening.
So there can be controls in that direction as well.
I mean, I live in Europe and not wanting to depend on external parties, be they North American or Asian, for critical infrastructure is a really hot topic in Europe as well.
Yeah, yeah, yeah, for sure.
So I think that's going to be an element of it.
And you've almost, we're almost into like globalism now.
And, you know, should, you know, should countries depend on one another or should we be isolationist?
Like we're definitely trying to solve the world's problems here, Chris.
Yeah.
Yeah.
We might need a third podcast to completely finish.
We might.
We might.
My only counterpoint to your argument around Jevons paradox, and we always want more of a thing.
I mean, yes, you're right.
And I could see it going that way.
Like it certainly has gone that way with save.
For instance, my first cool cable modem was 960 kilobit per second.
That's almost a megabyte, Chris.
Oh, my God.
A T1 is a $3,000 a month thing, and this is almost that, and it only cost me 150.
And now I've got a gigabit connection, and if someone offered me a 10 gigabit, I would take it.
Would I know what I'm doing with the other nine gigabits?
I don't know, but I would take it, right?
So maybe it goes that way.
On some level, it's gone that way because it was commoditized and because it was cheap and because we could keep saying yes at roughly the same hundred bucks a month.
I think the argument on the fax side was with facsimiles, it didn't really matter.
We didn't need to be able to send 7,000 page faxes in three seconds.
That was never of value.
that basically what the fax was doing was like, look, I got a wet signature, you know?
Well, it isn't wet, but I got a signature, you know?
And that's mostly what it was doing in the end.
You could send the other 25 pages in a Word doc, in an email eventually, you know?
And so good enough had a different bound.
And because, and I think prices were something about what kept that in check as well.
Yeah, yeah.
If we wanted to send that 7,000-page document in two seconds, we eventually did that on our gigabit line computers.
Because, again, that's where the commodity was.
That's where you could commodify and bring the price down.
So, yes, Jevon's paradox.
Yes.
I mean, I'm a big electrification fan.
And there is definitely a threat that the more we electrify, the more we use electricity, which might just mean we have to bring the coal plants back on.
And that's not what I'm after.
So, I get it.
I'm with you.
But I think there's another side of it, though.
As long as we can keep optimizing the hardware and we can stay just ahead of the efficiency curve and we can commodify, commodify, well, eventually I will have that 10 gigabit connection to my house and it'll only be $100 a month.
They probably already have it in Korea.
Yeah, yeah.
I think you're definitely right to think about diminishing marginal returns and what people will actually want.
I think that that's true.
And no company is kind of owed a living by the world.
So, you know, you make a good point.
And there may be people certainly part of the supply chain that are routed around like FedEx was.
Yeah, yeah, yeah.
It's interesting times.
Maybe the thing that I could add on to that idea about to what degree people will be willing to pay more to stuff.
I think I believe that that attribute of humanity might make it very difficult to improve the quality of software using coding agents.
So I think I do believe that if you were to.
do a great job with all the different kind of harnesses, the feed forward, the feedback, guardrails and whatever, that in theory, someone who then maintains the same amount of energy and discipline can do a better job than if they were coding by hand in many cases.
So I don't subscribe to kind of the slop narrative that implies that coding agent will automatically do a worse job.
I think you can do a better job with coding agent, especially if, for example, you invest some of the time you save in improving the harness or doing more review.
The conclusion I've reached, though, is that if we looked like we were accidentally creating greater quality, we would respond not by achieving that quality, but by lowering our standards or increasing our haste until we managed to lower quality back to where it was before.
Because the level of quality we have in software is not the maximum we know how to achieve, or at least I don't think it's been in any code base I've been part of.
It basically lowers.
the conditions will accept.
And if it gets lower and things start breaking, then people will get it so it's just above the waterline.
So there's a paradox called Worth's, it's Worth's Law or Worth's Paradox after the computer science.
And that explains why using a computer doesn't get faster.
Like you think it should get faster to use a computer because computers are getting faster.
How could that not get 12?
Because the faster the computer gets, the more layers of crap we...
put on our stack until it slows just at the point where we can bear it.
You would think, you know, I've heard people say many times, well, hang on a second, can we just stop developing more weird layers of abstraction for like one generation of hardware and just take the win?
Wouldn't everyone love that?
And maybe they would, but certainly the revealed preference is that they don't really give a crap of that and that they're willing, at least us as software developers are willing to fritter away.
So I think...
We might have the same quality equilibrium problem where even if the techniques are available to produce high quality software with better modularity, with more stability, with fewer bugs, with toting agents.
If we achieve that, we'll say, well, that means we can take a little bit less care next time.
That means that we can go a little bit faster next time until the quality drops to the level that people were kind of...
grudgingly accepting anyway.
So I think there's an equilibrium force there.
Oh, I like it.
If I borrow the analogy for a minute and riff on it, what's coming to mind for me, and I've just now heard of Wirth's law, so I'm probably going to get it wrong, but I think it is sort of a counterpoint, I guess, to Moore's law.
And I always thought, like, yes, you know, CPUs get faster and faster and faster, so why don't they get slower?
And I always figured it was sort of two debts in the budget.
Two things that that was going to, maybe three.
Like maybe one is like, okay, you know, software can go faster, so we don't have to performance optimize as much.
So that's like going to the margins of the software development companies in theory.
They didn't have to spend as much money on performance optimization.
So I think that's one bucket that it's going to.
I think another bucket that it's going to is like shiny features.
Like you and I right now are in different countries of the world, and yet we're having a very high resolution conversation as though we were not.
That was not possible on my 4D6 DX33.
It just was not.
And so that's like a budget.
But I think there's this other budget, which is like technical architectural debt.
Like Linux still has POSIX in it.
And so does Windows.
Because for a minute, we all thought that we needed this layer of translation called POSIX, which from what I can tell, didn't ever do anything, right?
Like no one ever actually got any value out of POSIX.
I'm probably wrong.
Probably someone did.
But all the computers still have a positive layer in them.
So I figured, like, if Moore's Law ever did stop, that it would, hopefully the budget would come out of that.
Like, we would sort of go back to our technical debt, these old abstractions that don't matter anymore.
And, like, let's strip them out of the operating system.
Are we going to get a couple of milliseconds back?
I don't know.
So how does that apply to this?
I'm not sure.
But what do you think?
Well, I think, well, I think it...
It's a neat example because it works for both technical debt for quality and for speed.
I guess I think we'll find our ways to screw up software in different and imaginative ways in the future is maybe my conclusion.
So it will be possible, well, it's always possible, possible with CodingAgent.
It will be possible to maybe refactor or remove code where we have a clear idea of what the outcome should be.
But the human labor to do it would be...
but we can specify the problem and put a harness around it really clearly, then we can go ahead and do it.
Still, you have to worry about compatibility and distribution and things like that that would be difficult if people are using POSIX compatibility as something as the interface that wouldn't be as easy.
So it's not that the same specific subsystems that slow us down or exactly the same anti-patterns that cause quality problems.
will be present.
So maybe in the future, C-based memory bugs won't exist anymore because we'll use coding agents to rewrite everything in Rust and that'll be done.
But I think there'll be a substitution of other quality problems that we will tolerate up until our level of acceptability of our employers slash the patients of our users.
And just at the point before they throw the laptop down in disgust, we will step in and just Just keep it just good enough.
So I know this is kind of a pessimistic take, but I'm just trying to imagine the organizations I've seen tolerating software improving.
And I just can't say it.
If there was a bunch of engineers that came in and said, okay, great news, software is twice as reliable, let's say.
I'm talking about kind of bugs and things as it's been before.
A lot of the responsive organizations will be, well, why did you spend so much time on it then?
Unless maybe as consumers we can demand our share of the surplus introduced by coding agents to be paid to us and increase stability or performance, I think it will get eaten maybe by many other things into the profits of the company that is sitting at the right point in the value chain.
Yeah, 100%.
I can remember, and this is now 20 years plus, this is pre-ThoughtWorks days for me, so I won't name anybody, but it's nobody in my life currently.
Going to the CEO, and this is when I was first discovering test-driven development and unit testing, all of this stuff.
And I remember going to the CEO and saying, boss, we're going to start unit testing everything.
Quality is going to get better.
It might slow things down a little bit.
I don't know.
And he's like, I'm not sure that quality matters that much, Kyle.
No word of the lie.
And he's not wrong.
You know, the investor put money in because they want to see money out.
He wasn't wrong.
Now, luckily, the unit test sped us up.
It didn't slow us down.
because, you know, it just helped cut things out before they launched.
There is the quality is free, which I think definitely does apply to a certain kind of quality.
And unit testing is great.
Definitely a quality is free phenomenon.
Yeah, yeah, yeah, yeah.
But not all of them are.
The other thing that's coming up for me is they originally, like internet lines were sort of like the same telco lines.
And telco lines were built for telco levels of reliability that needed to be there because they were centrally managed and circuit-switched instead of packet-switched.
And people started just saying, well, screw it.
Just make it an Ethernet line between two cities.
And it obviously can't be the same copper coax.
You'd end up having RF issues and so on.
But still, drop the telco levels of reliability that telco expected because the Internet was architected not to care as much.
about reliability.
And that, that cause, I can remember buying a T1 line in Kingston, Ontario, and it was $3,000 a month.
And they, you know, I got a quote for, you know, for basically the same meg and a half from somebody else.
And it was like, I don't know, $300 a month.
And I remember calling up my bell rep and he was like, well, you could, but the quality is going to be terrible.
But Chris, it wasn't like, it was fine.
You know, like maybe the likelihood that, that there would be a problem.
Well, that's not even true because later, I don't know if it's the same year, but a bell technician dropped a wrench in an open router and took out like most of Toronto for a couple of hours.
So even my $3,000 T1 line.
And that's when I sort of gave up and said, okay, this is dumb.
So I do think there is something like that going on around quality where there's this sort of like, if we have to just understand now that.
Everybody's going to be doing four things at once.
Everybody's got an app that they just put in the marketplace.
Everybody's got a startup that they launched.
We're going to have more things, right?
Did as much care go into each one of them?
Probably not because you could do it faster, which is why you did it in the first place.
But is that a bad thing?
Maybe not, right?
Like it might be like that $150 T1 line where it really was okay in the end.
Yeah, some of the quality...
Maybe it actually serves humans' needs better than if it was...
high quality in a purest way, that could be true.
Yeah.
Yeah.
And that's not to say it'll only go well.
I'm sure in places it will go terribly and we'll tell interesting stories and it'll be funny.
Yes.
Well, I guess, you know, with all sorts of feedback loops, if it takes a long time to figure out that you've done something bad, then you're much less likely to gracefully adjust, you know, a steering wheel of the car where you can.
where you can correct, course correct instantly is fine.
If you were trying to drive a car with a steering wheel that had a lag of, I don't know, let's say a second, your chances of getting through, you know, seconds is not very long, but your chances of getting through traffic unscathed would be very low.
And so the analogy I'm making is that for things like, I don't know, accidentally lowering the quality of the software in your estate or poisoning the well of open source, for example, if it takes...
A while before the implications of our actions to come back and bite us, we're not necessarily going to take holistic, long-term, rational decisions.
We might find that we've messed things up before we know it.
I think it's generally thought of that we are bad at seeing those kinds of things as a species, in fact.
Yeah, well, maybe if I could relate it to my theme.
I think we talk a lot.
in terms of feed forward about how we respond in a society.
Oh, we should do this.
Okay, everybody, let's do this.
We imagine that as how as human beings we make decisions.
But I think feedback and things going wrong and being forced to do things when there's no alternative is much more how things change at a social level over time.
And so if we're in a context where we don't have good feedback, and maybe this relates to it.
the climate change question that I said before.
If our actions and our consequences are separated by enough distance, it might take us a while to change what we do.
Yeah.
Fingers crossed.
We're going to need all the luck we can get right now.
I think so.
Chris, thank you so much.
This has been great having you on a second time.
Thank you.
I feel like you got pessimistic, Chris, this time around in various ways.
I've definitely really enjoyed the conversation and hopefully we've increased the quality of our ignorance together.
Amen.
