# AI Software Engineering: Production, Trust, and Platform Strategy

**Podcast:** Thoughtworks Technology Podcast
**Published:** 2026-07-09

## Transcript

Hello, everybody, and welcome to another edition of the ThoughtWorks Technology Podcast.
My name is Ken McGrage, and I'm very honored to be met with a couple of our ThoughtWorkers from our UK offices.
I'll let them introduce themselves.
Andrew, do you want to go first?
Yeah, hi, Ken.
Thanks for having me.
I'm Andrew Harmer-Law.
I'm a tech director, like you said, based out of the London office.
Great, and Keefe.
Yeah, I'm Keith Morris, distinguished engineer, wrote the book Infrastructure as Code, so I'm kind of largely known as that and also based out of the London office.
So part of the reasons that we have two of our folks from the European office is that ThoughtWorks recently hosted an event in Europe last week around the future of software engineering.
And we were really privileged to meet with, gosh, I guess about 80 folks.
from really around the world, but mostly from Europe, that are real leaders talking about the future of software engineering.
Now, it's easy for me to say that, leaders, but I will have to let you know right up front that we did run the event under what's called the Chatham House Rule, which that rule states that we cannot use the names of any participants or their organizations.
This gives people the ability to speak freely without worries about misquoted.
you know, to give their personal opinions and so forth.
It's really important to an open discussion.
I encourage you to check it out.
It does make it a little hard because we're going to tell you that these were knowledgeable people and you can either believe us or not.
But we did a similar event a few months ago.
It was interesting.
I was privileged to be able to go to both.
And just a couple observations and then we'll get into some folks and things.
You know, first off, it was observed by somebody else that in the previous event, which was in the United States, it was very much about, oh, look at this cool, shiny stuff.
I mean, it was this year, right?
So these things were not brand new.
But there was still a lot of, oh, we think this might work and this might work and so forth.
And in this event, it was a little bit more practical of, okay, we know this can work, but we need to answer these questions.
So, I mean, part of that is cultural between two different areas, of course.
Part of that is just a different group of people.
There's probably some observation bias, if I'm honest.
But some interesting takeaways to get into here.
And so, first off, for those of us, and I know this is audio only, but you're listening to three...
white people with gray hair.
Keith has less gray hair than Andrew and I, who've been around a while.
And, you know, we heard people calling like 30-year-old Cobol legendary coat.
You know, it's like, okay, you know, give me a little bit of a break here.
Anyhow.
And so it's interesting how fast things are moving.
And so I guess first off, just, you know, if you don't mind, I'll go to each one of you and just some of your impressions overall of the event, any major takeaways.
We'll go through some individual questions, but I'd just like to get your gut feel.
And I started with Andrew before, so Keith, I'll start with you.
Yeah.
So that comment was interesting.
So I did talk to a few other people like yourself who were at the event in February.
And I wonder if part of that difference also was the timing, because February was the point where there had been the new models dropped in December.
And that point has been kind of identified as a turning point where like suddenly the models got really good at coding and it became a lot less about like trying to, I don't know, squeeze something usable out of them and a lot more about, oh, now it's producing stuff.
And so I think that was kind of fresh.
Now we've had, that was what, six months, five, six months ago.
And now we've been doing it for a while.
And I think now the kind of focus is on, oh, well, we've kind of learned a lot about what challenges this creates, you know, having models that can produce good code, but yet there are a lot of challenges in how they do that and how we make them work the way we need to work at the results we want to get.
And so I feel like a lot of the conversation was focused around the challenges that we've kind of discovered and are trying to figure out.
Andrew, your impressions.
Plus one to everything that Keith said.
It's definitely things that struck me was because now it seems more inevitable this time that this stuff is useful and we can't pretend as software professionals that it isn't useful.
And now there were lots of themes, but one of the big ones that stood out to me was the fact that people were like, right, how much, where do I sit in this whole thing now?
Because we know the typing went away, right?
Like there's less typing, but it's like, how much do I still need to bring to the party?
And like, how much do I need to, like a good phrasing that I heard was that, how much do I need to trust this stuff?
Like some people were like, you should trust it completely.
Other people were like, cause you can get, you can build 17 different copies and then just see what they're all doing.
Or other people were like, no, you need to look and inspect it and all this kind of stuff and be very, very careful about what it does.
And so the lots of discussions were about that.
And then that kind of fed back to another thing that's making me think a lot.
is like we evolved ways of building and running and kind of operating and evolving software in a world where we created the software by hand effectively like bespoke and now other things are potentially doing that with us as partners or maybe like like dark factories got mentioned as well right so we'll talk about that like and now we're like where do i sit in this and which skills to have do i do i still have this in this and where where are the leverage points for for me and all of this kind of stuff and but it wasn't like scared like we none of us have jobs anymore it was more like right where are we recalibrating that was loads of the conversation and then you get specific conversations about this yeah because i was having a discussion at lunch with somebody and he was saying this is a person that again had been around for a very long time and was saying yeah where before i would try to to describe six different ways of working now i just do six prototypes um And it was like, okay, but how big is that code based?
And what are you working on?
Because doing six prototypes is one prompt maybe, but it's some number of tokens, right?
But then I asked the, you know, the big question at the end, well, how do you actually run them or how do your customers run it?
And so I guess, you know, Keith, you know, we hear about AI for generating all these, and I'm just going to do, you know, seven different versions for user experience.
how do I run this stuff?
What does it do to the infrastructure's code calculus?
I mean, and I know there's not a clear differential between application and infrastructure, but, you know, where's that last mile come in?
Yeah, so we had like two sessions that come to mind.
And I can't remember if you mentioned that there was the open spaces session, so it was a lot of kind of general discussion rather than people presenting, right?
And so two of the sessions around that that came to mind was, so one, we talked about path to production, so pipelines and this kind of stuff, which...
to me it was about how do we prove the software is production ready um and i think that relates a lot to like so what are you using the software for a lot of a lot of cases people are building software is kind of like oh i'm gonna make a tool for myself to do a thing and that's kind of fine or like me and a couple of people my team are going to use this thing but when you're doing it with like software the company is going to use or that is your company product um it's a whole other level of you know from knocking something up with a you know, Replit or what have you.
And okay, I can, you know, I've got this working software to like, yeah, and this can scale to, you know, thousands or tens of thousands of hundreds of thousands of users, and it's going to be secure and it's going to be stable and all those good things.
And so I think the kind of the conversation we had in that one session was around, well, how do we use the pipelines to make sure that our code is going to be ready to deploy?
And we talked a lot about where do we fit agents into the pipeline?
Like how Far, like, do the agents get involved in what happens after you push to the pipeline, then automatically kind of take the results of, say, a failed pipeline stage and then make corrections and push another build through and so on.
Yeah, some of the stuff we discussed in that session was around like that gets expensive with the infrastructure as code.
You're deploying your AWS resources over and over, and the agents tend to kind of thrash a lot on some of that stuff.
And so that's something that like, I don't think anybody seemed to have that really solved.
I read a friend that works at Apple and he said, you know, if he's doing Python or whatever, he'll use agents and doesn't even look at the code.
But if he's using someone like Swift or whatever, okay, wait a minute.
They don't have as much training data.
You know, what's the maturity of some of these tools when it comes to, you know, I'm running Terraform, you know, versus Java.
We were talking about how much more effective they are in code.
Is that true on infrastructure's code as well?
Is there a big enough training set or does that still require more oversight?
That's a good question.
I mean, I don't think we talked in the sessions as much about that specifically, but my experience has been that the LLMs know a lot about infrastructure and they can talk about it a lot.
But when it comes to like building code, generating infrastructure code and then deploying it, Like they make a lot of silly mistakes.
I don't think they're as strong maybe as they are with like writing UI code or just kind of code that can be run and tested locally.
So I think that's why they kind of thrash a lot more is like, you know, generate some code and then it, you know, they try to deploy it and it doesn't work.
And so they kind of go back and forth and like, how does this actually work?
So, yeah, it's, I think they've got a ways to come.
And one of the things we did talk about was around using, so these things where there are kind of like shortcomings in how the agents handle stuff like we talked about in the path of production.
We also talked about like in operations and kind of diagnosing runtime stuff and those sorts of things that providing more forward kind of guidance, like skills and so on that are a bit more directive to like, okay, here's how to do this.
may be necessary to get better results in those cases.
When you take skills, are you talking about specifically like...
I'm talking about age and skills.
Yeah, okay.
Yeah.
So that's a person's ability.
Well, I think there's a combination.
This is like, you know, we talked as well in sessions over all about harnesses and all this kind of stuff.
And as Andrew mentioned, the trust, how do you get to trust the code?
So a lot of the conversation was around what are good ways to do that?
And so skills, as in the markdown files that you give to agents to say, here's how to carry out a task.
And, you know, the agents MD, Cloud MD and specifications is one thing, but that actually tools like scriptable tools that are executed and do particular tasks the way, you know, in a predictable way are a bit of a stronger.
So there's kind of like levels of strength of assurance that you're getting these things to do things in the way that they should.
So, you know, the loaded question, you know, that I don't really expect an answer for, are agents members of our teams?
So I guess, Andrew, you know, you talk a lot about team structure and, you know, Conway's Law and myself are about the same age, I hate to admit.
You know, where does that come into play?
I mean, if, you know, if Conway's Law, if our code, well, I'll let you tell people first, what is Conway's Law?
So Conway's law is that the communication structure of your organization, your software will follow the communication structure, not the org structure, but the communication structure of your organization.
Generally, there's this whole thing.
It's like if you have numerous teams and they communicate in certain ways, then you'll see that kind of.
So it's like if you have three teams building a compiler, then you'll get a three pass compiler.
So then what's that look like when it's one human and 73 agents?
So I was paying really close attention to this because I'm.
As people might know, I'm a DDD nerd and I love bounded context.
And I love, there were sessions about modularity and like, like we never, modularity never really worked.
You know, it's not like modularity was clean before agents got involved and now it's getting all messed up.
So there were sessions about that.
And like, how do we make things, you know, is this an opportunity to maybe make things better?
And like, I've been doing experiments that I talked about a little bit in a few sessions, but like, I've been trying to figure out how.
how much of domain driven design agents know and stuff.
And then meanwhile, everyone just seems to be, like Keith said, we're coming out of the phase where we've all realized this, they can, the agents and, you know, the LLMs can do some very interesting stuff and like stuff that we didn't think they could do.
And so a lot of people were just throwing requirements at them and then generating code and stuff.
And as long as you didn't look at it too closely, it did functionally what it was supposed to do.
Maybe non-functionally, it was a little bit, you know, the edges were a little bit blurred.
And so we weren't so much worrying about the structure.
And the things I've seen, and I heard this from a few other people as well, is like, because agents can kind of roam around anyway, and they're very sycophantic, right?
They want to do what you've asked them to do and make you happy.
So if you want them to do that while still kind of discovering and evolving and protecting some boundaries in a code base, they...
highly possibly will not respect that, right?
And so even if you start with something clear, the edges can get blurred very fast and you might have to wait and then tell them to unblur and then they'll get blurred again.
And then so lots of discussions we were having was like, is it time to go back to individual microservices in separate repos purely because you can give an agent, one agent, you can say you can access this repo, but you can't access this repo.
Kind of to Keith's point about like different levels of kind of firmness that you can put in things because sometimes you need to be quite firm and quite predictable like this.
This agent can't change this because it has no rights to change this code base in this repo.
But then, and this again was another kind of conversation that spilled over into the hallways afterwards, if the boundaries were easy and obvious to find in the first place, we'd be good at it, but we're not good at it, right?
Because these boundaries are quite difficult to find and then they evolve over time.
And if it was easy, there wouldn't be the entire blue book of domain design from Eric Evans.
So all of this kind of stuff, it feels like...
to me anyway, that there's like another layer, kind of like what Conway's saying, right?
Somewhere on the interface between teams and the software that they build where we're putting this into, we're encoding this explicitly, maybe into our agent swarms or agent pools or whatever, and the teams that are controlling them to make, to get this kind of architecture and maybe inverse Conway or whatever it's called, to kind of hopefully drive these things in a sensible direction.
But at the moment, it's, there's a lot of people been like, Keith said, kind of everyone's been building a lot of stuff.
For either one of you, what do you think about this concept?
One of the things that came up with sessions was giving people, and I use this term because this is the term that was used in this session.
I personally am not a fan of using this applied to agents, but they said, you know, they gave their agents personalities in that this agent is a cynic and this one's an optimist and this one, you know, what have you.
And what do you think about that overall concept?
Should agents have different, ways of looking at the problem?
I've got a psychology degree.
So I'm very human beings.
We're very good at anthropomorphizing stuff.
And it already worries me how much we're anthropomorphizing stuff.
So I wouldn't give them personalities.
You could give them hats, right?
Like there's the Edward de Bono black hat, green hat.
You could make them.
And I've been messing around with like different agents being different domain experts, right?
So you'd want a domain expert to be like, this boundary context is mine.
So therefore I understand it.
And I'm going to watch out for other.
other agents kind of straying into my ubiquitous language etc etc i think that makes sense i would worry about like like you said the terminology and you could see people me included we'd say things like reasoning and then we'd be like well or like you know we'd be kind of backing away from some of the potentially not entirely proven language about some of the powers of what the lms can do i'm really uncomfortable with anthropomorphizing for a number of reasons obviously there's kind of the psychological things you know that that people get into There's also just that whole kind of idea, you know, when you have like a business leaders talking about, you know, I'm going to have agents rather than people and all that.
And that's just wrong in a number of ways.
I think it's practically wrong, even if you're kind of like want to be cynical and, you know, focused on the, you know, I don't know, the bottom line or what have you.
I still think it's not a good framing, a helpful framing.
I had a.
interesting conversation and this was kind of like a follow-up um afterwards where i was talking with a couple of people and they and uh you know from the from the conference and one of them was talking about like oh i want to make more use of agents like i kind of my eyes have been opened a bit to like what people are doing with these kind of teams of agents and all that um and so i was reflecting with them on because they were talking about oh i have like a ba agent and a qa agent and so on and i was reflecting that like i kind of started doing that at one point And that was kind of useful, almost, I would say, like training wheels of like as a way to kind of get started with agents.
But where I quickly came to and I think where, you know, so talking with people in a lot of the sessions, we had a lot of talks about workflows and those kind of things.
And I think where it comes down to for me is it's like thinking about those workflows and which parts do you want to hand off to an LLM to carry out?
And then what are the handovers between the humans?
Okay, I'm going to like.
whether it's a prompt or whatever, I'm going to indicate an agent, I want you to do this.
And then the agent comes back and says, I've done this.
And then I have a look.
Am I happy?
How do I feed it back?
And so on.
But it's very task oriented.
And I think that's a more useful framing as we think about what tasks you need to get done and what, you know, and so what are the agents doing?
And so having an agent where you've kind of like scoped it and saying like, okay, this is the task that you're going to carry out.
Here are skills and tools for you to use.
And I think.
attitude in terms of like so you might want something that is thinking okay take an adversarial approach a skeptical approach sure i think you can do that in a way that's pragmatic without anthropomorphizing it a lot came up that in sessions i was in about the you know platform teams and um i doubt very many of our listeners have seen these people when i do i rant about devops team and how it shouldn't exist uh and that sort of thing But what is the role of, is a platform team just infrastructure?
I mean, and I air quotes with just infrastructure.
Who owns the harness, the organizational knowledge, that kind of stuff?
Any thoughts on those?
Either one of you?
You touched on it, Ken.
It's like, this is the conversation we've been having for a long time about, okay, once you get past having a single team that's building software, ideally running software, Once you get to it, you get to multiple teams, a lot of stuff going on.
How do you share?
How do you cooperate, collaborate?
And it's always a tough balance, right?
Because you can say, oh, okay, we're not going to have platform teams.
Every team is going to do their own thing.
And then you get a lot of divergence and not much alignment.
And then I can potentially a lot of friction from teams doing things in very different ways.
Like I've seen one organization.
where they had an e-commerce flow, they divided up their teams like different parts of the e-commerce flow.
You said one did kind of like the search, one did the product catalog, one did the shopping basket, one did the checkout.
And they were all each doing things like with their own tech stack, even like the UI used like different JavaScript frameworks across each of these, right?
It's like, okay, clearly we need some level of kind of coordination and collaboration and also learning, right?
And I think especially right now, because we're still learning how to do this.
And I think there's opportunities for us to kind of share.
and learn.
So if we're talking about things like harnesses, how do you build a good harness?
None of us knows.
So rather than everybody trying to build their own thing, what can we do?
But all of that, with that trade-off of that, if you have a separate team that says, we're going to own that and you're going to use what we create for you, that tends not to work out very well.
And I'll leave it there and pass on the batons.
Andrew, you've spoken a lot about decision-making authority.
Where does that fall in this?
Exactly.
So I think...
the decision-making thing and who gets to decide i mean famously i forgot to mention my book at the start um because i need more training on that but um so my book is about facilitating software architecture and like fundamentally like devolve decentralizing decision making and i write this book and an ai comes along and basically just decentral decision making has got hyper decentralized right to the extent that maybe a bunch of decisions are being made by machines that we didn't even intend to make decisions um So there's that.
The other thing that came up a few times, which was super interesting to me was, and where platforms came up again, and I think non-functional requirements come into this as well, as someone who tries to wear the architecture hat.
What's new about this, one of the big things that's new about LLMs is the exec can like vibe code stuff and then they want to deploy it and run it.
So Keith kind of mentioned this, alluded to this a little bit earlier, right?
So now the suite of the kind of empowered people, and you mentioned this a little bit, Ken, I think as well, like, There's more people who can write this stuff now, right?
And we're like, right, it was bad in the old days when everyone just YOLO'd things into production.
But now, like, the variety of things that could be YOLO'd into prod could be even broader than ever.
A, because it, you know, reduces things and reduces cost and makes things more predictable.
That feels like a good kind of target kind of thing for platforms.
But also, like...
If there are certain things, like if you're in a regulated environment, lots of attendees were, they're like, there are certain things we just can't do, right?
You need to be GDP, like everyone needs to be GDPR compliant, right?
But then in certain industries, they need to have certain like PCI compliance and stuff.
And so you could lecture people on it, or you could build a platform which goes, if you use these things and include up to and including harnesses and skills and whatever, then we have blessed these because we think they won't just give you the functionality, you prompt the functionality, but we're...
pretty confident that they will give the non-functionals as well right we'll store it in a certain type of database we'll just not store certain types of data we'll replicate it in some kind of way and then it strays into some of the conversations that keith mentioned around about like building these things is super easy but running them might have just got more complicated so again like operability observability and all this kind of stuff so i think it's very interesting and that's to go back to your question about decisions I think the key thing now is lots of decisions.
It's in my book.
People used to make decisions without even realizing it to begin with.
Now, a lot more people are making a lot more decisions that could be a lot more significant without thinking about it because it didn't confront them, right?
I think surfacing that kind of stuff, to Keith's point, it's like giving people tasks, maintaining responsibility for the decisions that you should be responsible for, et cetera, et cetera, I think is going to become key.
And I think a lot of that, was surfacing in people's thinking.
They're like, I don't want to know every detail about every single thing that every agent is doing, but I do need to have control over the things which I'm responsible for because I can't blame a robot, et cetera, et cetera.
And that might have come across as pessimism.
I would classify it more as like European realism than anything else.
But yeah, people are like, right, this thing is definitely magical.
Now, how do we like harness the magic for the forces of good and not chaos?
It's funny you use the word control in there.
And I am unashamed of the fact that I do use tools like Claude and that stuff to summarize.
Like I used summaries of, what, 40 different sessions for this discussion.
And one of the things that showed up was there is a recurring, quote unquote, theater of control critique.
You know, and so I guess, Keith, you know, what are your thoughts there?
I mean, do we actually have control or is it just theater?
So it's really, really interesting.
This came up in a couple of different contexts.
This idea of how much autonomy do you give to the agents, right?
How much do you just let them get on with something?
And what do you have to check?
And so like a couple of the contexts, there's obviously the coding one, it generates code.
Okay, do you have to inspect every line of code?
And there is that kind of like, There are some people who are very much in the camp of, yes, you do.
And there are those in the camp where it's like, oh, you don't have to anymore as long as it gives you what you want.
And then there's a whole bunch of questions around, well, how do you validate that it gives you what you want?
Which is where the harness stuff comes in, which is also talked about in the context of building knowledge graph.
So there was one really interesting session where we talked about building a knowledge graph of your organization, your domain, and your business and making that available to agents.
um and a kind of a rag kind of a way to where it can kind of you know draw the right information that it needs without having to fill its context with everything um so how do you build that knowledge graph is that something that humans build or actually then we were talking about well you can get agents to do it but then how do you know that they're doing the right thing and so putting that onto it and then the third kind of context was when we were talking about the operations the day two um what happens when you have stuff in production and can use this stuff for troubleshooting and again there's a real divide between those who are like it's fine to use them use agents to kind of gather information, present you with options and so on.
And even some people were a little bit uncomfortable with like having it suggest here's courses of action you can take because that might then bias you and you might not think about other courses of action, right?
You might trust that too much.
And then the people who are like, oh, just let it go ahead and fix things, right?
And so that was just really, really interesting of like, and so do we have control?
And I think a lot of this comes down to, did we ever have control?
So there was, I think, probably what surfaced, Ken.
Yeah, yeah.
Ken, when you were talking about what surfaced and kind of like looking across the sessions, I think a number of times the whole thing of pull requests was brought up around where you have people out there saying, oh, you have to have pull requests to manage your agents because otherwise you can't.
It's like you can't catch all the things that they might do wrong.
And it's like, well, are pull requests ever a guaranteed way to catch that like a human is doing something wrong in the code, much less an agent?
or is that just theater has it always been just theater um and so that was like a really interesting uh observation and andrew i'd ask you to do the same question because i know you have an interesting outlook on this i mean it's literally the first chapter of my book like i would argue we've always had a lot less control than we thought we did um and my hypothesis in the i mean like again it's it's interesting like the now I'm not sure I agree with everything I wrote in the book now.
I agree with most of the stuff I wrote in the book, but like the framing of some stuff, because I'm like, crikey, now I'm being confronted with even more anarchy than I thought I was advocating for.
But we always had less than we thought, right?
There's an Alberto Brandolini quote, which I love, which is one of the three things that kind of informs the book.
And it was like, it's not the design that goes into production, it's the developer's assumption.
Now, the developer, the person who wrote the code, might be an ALLM.
Right.
So like, and it'll have assumptions because it's, it's a large language model, right.
It knows a lot about all those stuff.
Right.
Um, so that's interesting.
So it was always probably impossible to guarantee that the, you know, the diagram you had of your domain model on Confluence was the thing that was actually running in prod.
Cause we all knew it wasn't right, but some people could kind of kid themselves.
Now it definitely isn't.
And like figuring out which bits, which bits we care about most and which bits we kind of.
we need we think we need to intervene in and which things we need to be like some some other folks were like and i think i agree with this like a crusade against the word determinism which seems to have caught on but like predictable right which things we need to be predictable where we need to invest confidence can i challenge something you were just talking about the diagrams right as an example of the architecture diagrams and confluence like we know they're never up to date and all that isn't there an opportunity now where it's like you know we're okay we're too lazy to update it but can you just like have the agents update the diagram update their diagram and just like and so it could actually be more up to date than it was before and that could be the level of engagement right this is the thing this is why i'm still and i'm literally i'm still kind of this is about beautiful open spaces but then it's spent takes six months to process everything we totally could right like that the knowledge like you were saying the knowledge graph and basically like the bigger the knowledge engineering thing i think that's might be a term that we end up using more because we can't not have the knowledge we can't someone needs to be doing the thinking somewhere maybe with a lot of assistance maybe like we're devolving lots of details to other things but that feels more important right and how we go about it i think is not clear like i was in a session where we were like how much we want to hold on to practices like tester and development and different people have different opinions um But the reasons why we do TDD, for me at least, to kind of break open the process of generating code into the thinking about what it does, thinking about how it doesn't, thinking about the design of the implementation, those three steps are still relevant.
And we care about them more or less in certain degrees.
So all of this kind of stuff, pulling that to the right level, I think would be very interesting.
There's another point I wanted to kind of capture on this, which is...
A lot of these conversations we have about the control and what do you trust the LLM to do?
What does a human have to monitor and all that?
I think one of the things that we fell into a lot was this idea of thinking that it's static.
Like the answer is this right now.
Whereas actually I think it's like, so like in the, in terms of like the debugging and troubleshooting production systems, right?
I think there's probably different classes of problems where there are some things where it's like, okay, you could have the LLM.
Especially if you're not, again, if you're not just kind of having like, say, Claude, look at the system, watch it and fix anything that breaks, right?
Just do whatever you think you ought to do.
But more specifically, like, okay, here's an alert for a specific thing that happened, trigger an action.
That might be an LLM making some decisions on there because it requires a bit of flexibility and stuff that you can't kind of predict as much ahead of time.
And so there might be some very narrow scope things you could do today and give it that kind of latitude.
And then over time as we learn this stuff, and so the stuff you were just talking about, Andrew, of like, you know, all the kind of governance stuff and the information gathering and all that kind of stuff, we'll get better and better at that and be able to provide more focused and more confident.
And so I think we'll increase our maturity as an industry and what we can trust the agent to handle autonomously.
Yeah, I tried to float this, but I don't think it stuck.
There's a benefit of, again, of open spaces.
I think there's a concept of surprise, right?
Like that, because again, my degree a long time ago was in cognitive psychology and stuff.
Control is one thing, right?
And we definitely can't be in all the right places to do control.
But we want to control things more where we want to be surprised less, right?
There might be some stuff where we don't care, right?
So I don't care.
If you've used one of 17 different JavaScript frameworks, but I don't care, then that's not very surprising because someone would use it.
But if I'm like, right, this probably needs to meet these regulatory requirements.
And if I find out that it is writing PII all over the logs, I'll be surprised and it will be a bad thing.
And like, I think that's something that LLMs can watch out for.
Things like embodied systems, which we kind of didn't really get onto because they're not LLMs.
That's kind of one thing that some models of that are tending towards, right?
Like watching the running systems and keeping an up-to-date mental model in some form or other, like a predictive model of what the system is doing and should be doing, and then looking for anomalies and all that kind of stuff.
feels like it's still way off but it might not be as far off i mean i keep thinking this stuff's years away and then it happens three days later okay so so i recorded a podcast some number of months ago with nathan harvey who who was not at this event so i'm not breaking the chatham house rule um and they were talking about ai amplifying whatever's there whether it be good or bad uh and the example he used was uh the the high school band that's Maybe not very good.
And if you put an amplifier in front of them, you don't get good music.
You just get louder, bad music.
And so I guess I would ask you folks, you know, I mean, we've trusted untrusted sources before.
There's ways to manage this.
You know, I'll start with Andrew, I guess, you know, Stack Overflow is the things, right?
I mean, now it's AI read and not us.
But what are your thoughts about the amplifying what we have?
My take on this is that if we understood things were important and we did certain things consciously, then we probably understand that they're a big deal.
And so they're important to us.
And I think if we didn't, we just kind of like just followed that, you know, if we cargo cult, we might have looked like we were doing the same thing, but we just were cargo culting it.
Like we were doing test room development because I ended up with a good test coverage or did test room development to think my way through a problem space, right?
It might look very similar, like the type things on the screen, but the motivations might be different.
And I think that you definitely could kind of end up going off in those different directions.
I think the intent is still a really big thing.
What I do think, however, is like, like Keith said, right?
Very smart people were saying, I just give some really good prompts.
to some LLMs and then get 17 different versions of the system.
And, you know, it wasn't like if I know a lot about this practice, I'm going to be hyper-controlled.
And if I don't, I'm going to YOLO stuff.
That was the interesting dynamic to me.
Some people were like, right, this stuff, I can suddenly benefit from, you know, multi-threading myself and doing all of these things.
Seeing how different people are relating them and their skills to the tools and the possibilities of all of these tools is very, I think it's still very much up in the air, right?
I think lots of people are still trying to figure out this stuff.
Things that did stick, however, seem to be no one debated like we move in smaller steps because it reduces risk.
You want to save code to like resource control.
At some points, you've got some kind of ground truth that you can come back to, et cetera, et cetera.
You need to have observability on your systems and all these kind of things.
You need to communicate and have kind of appropriate knowledge boundaries and things.
But other stuff seem to be very different.
People have very different viewpoints from my perspective.
Yeah.
On the amplification thing, I think, One of the things that it tends to, I find it tends to expose if your team builds software and you've got kind of those gaps.
So those things that like, yeah, we have the intentions to do, you know, we feel like we should do TDD and have good kind of unit test coverage or what have you, or have good unit tests, actually meaningful unit tests.
But we kind of don't, we're not really where we want to be with that.
We have some kind of like.
CI, CD-ish kind of stuff going on, but like there's still manual steps or things that we have to fix up after it gets deployed or something breaks.
And so I think there's a lot of stuff that we kind of used to get away with by just kind of like working around and all that.
Like we didn't necessarily have to have, or many teams didn't have to have the really hardcore, strong engineering habits and practices that they might've liked to have had because of the pressure to go fast and all that kind of stuff.
Well, now, AI spits out code so fast, you know, at a pace where it's just like you can't, those things fall apart, right?
You can't be constantly pumping things through a pipeline if it needs somebody to tidy up after every release, right?
Or there was a good example where someone talked about being worried that like an agent, if you don't review every line of code, the agent might slip something in where in the production environment, it tries to connect to the production database or in the development environment, it connects to the production database.
to kind of work around something or what have you to get data.
And gosh, that's why you need to have somebody review the code so it doesn't do that.
Well, no, you need to have infrastructure and systems that don't allow something running in your development environment to connect to your production database, right?
But it's like, again, a lot of times we just get away with that stuff.
We've kind of like gotten away with that kind of thing.
And now people use the analogy of like a plumbing system where it's like you're now pumping a lot more pressure into the pipes.
And so...
Things are going to start bursting here and there.
So I think for me, that's one of the kind of things that you see with that whole amplification that you get.
Yeah, stuff that might have been working by accident because you just weren't going at a certain speed.
Suddenly you are going at a certain speed and you're like, oh, we don't know how to change gear.
We're now trying to go 70 miles an hour in first gear.
And it turns out that the system doesn't like that, right?
Or the brakes don't work very well when you're going that fast, but they were fine when you were going 20.
Yeah, because you could break with your feet when it was going kind of slow, but now it turns out like you can, yeah, when you crash it.
I mean, to go back to the NFR thing, I think it just never fails to amaze me that the number of places I go where NFRs are kind of implicit and stuff, and I'm just like, this stuff feels very important.
And now I think more and more people are going to realize that just at least thinking about certain key ones for whatever they are, whatever is key for your system.
I mean, maybe I'm just being hopeful, but it feels like it's going to become bigger.
Are there any answers emerging?
That's a good question.
I think it's interesting.
Like it would be interesting to me because now that there feels like there's not like a marketplace, but there's more like people are sharing agents and skills and stuff.
And maybe to the point, like Keith for making like, how much do you care about the inside inner workings of your code?
Maybe like the, the, the old rising tide, tide raises, raises all boats.
Right.
Maybe just because lots more of this stuff is kind of becoming downloadable from like some skills database or something, then you can just drop it in and then you'll be magically working in a world where everything works.
I mean, I'm not sure.
I'm skeptical when I say it out loud, but maybe.
I think we have to get better testing that stuff, right?
And that's one of the kind of things, like when we talk about harnesses and harnesses and we talk about harnesses and harnesses, but like it's a thing, right?
And I think we're still not, when it comes to like skills and things like that, when I see some...
There are a few people there I know who have like really strong processes around using agents for the whole kind of like taking, you know, issues reported to an open source project and putting it through a pipeline to test and make sure it's okay.
And they use skills for that.
And they're using evals heavily on the skills.
So it's not just about bringing down stuff on the outside, especially those kind of agent pipe, you know, things and marketplaces and whatnot.
But like, how do you have something rigorous in place to kind of make sure it really does what you want?
It doesn't have any kind of security implications, injection attacks, those sorts of things.
So it's just, I think this is something that in terms of answers, I think one of it is it's like, we're learning to up our game and we have to continue to learn to up our game and stuff.
I used to work with someone who, their measure, they were a delivery manager and their measure of like a healthy project was how much you, number one, how much you'd hear teams talking to each other.
And number two, what they were talking about, right?
Because they could be talking about stuff which was panic-driven.
So that was bad.
Or they could not be talking to each other.
That was probably also bad as well.
Like there was no communication.
But like some healthy level of like seemingly important conversations between the important teams in the right kind of way, that's, I don't think that's going to go away, right?
Probably, maybe what they talk about is going to be different.
But you could probably, you know, you'd be like, okay, but this, why is no one worrying about PII or something?
like you said, like, you know, are separating production from development and so on.
In terms of wrapping up, I want to ask you both the same question, just your answer.
I'm just very curious because you're both very experienced in related but different places, right?
So going forward, at least for the foreseeable future, which I think in today's days, about 10 days, you know, right?
But, you know, not being glib in the foreseeable future.
What survives and what changes, right?
So what have we been doing that we should keep doing?
What do we need to start doing or stop doing?
I think, because again, this is, I wrote about this in my kind of reflections blog post.
The things I think are going to stay the same, but I think it'll become more explicit is like the thinking about the system, the thinking about, and at all of the different levels and being kind of conscious about the system and the structure of the system and what the system needs to do.
One of the big things I think that affected software architecture was product management came along and forced us to answer the question that Eric Evans asked, right?
It's like, what is the goal of the software?
It's not to scale Kubernetes in a horizontal way or have the right appropriate number of Kafka brokers.
It's like, understand this business problem and solve it.
I think it'll hopefully focus the conversations on that.
And maybe on how it gets down to implementation, but I think it's...
That feels more important now than ever, especially as we could be going faster and we could be building lots more smaller pieces that talk to each other.
I think that's really important.
What's going to go away?
I think, so the thinking I think is so important.
I think some of the practices that we've got used to to help us think might change, but I think that doing them will still survive in some way.
So I'm expecting to see maybe even more collaborative modeling, for example.
I'm expecting to see more awareness, like I said, of non-functional requirements, that kind of thing.
I now sound like I'm wishfully thinking.
But, yeah.
So I think the focus is going to largely be on the kind of slightly higher levels, the architectural kind of stuff, the stuff you were just talking about, Andrew, about NFRs, CFRs, that kind of stuff.
The conversations around that have to sharpen the things that we've had for a long time, but we need to get better at.
And I think what'll be interesting is like, I don't know about like the lower level stuff.
So there's a lot of things you talked about, you mentioned, and you mentioned the quality thing and like, do we care about what goes on inside the code and so on if we get what we want?
But I think defining what do we actually care about at that kind of higher level of abstraction and then understanding what do we want to guidance, do we want to provide to the agents to get that?
So a couple of examples, as you mentioned, there was a lot of talk about TDD and like, do we still have the agents do TDD?
Like, do we think that's a healthy thing to have them do or do we think it doesn't really matter anymore?
There's things like the code design, readable code and all that.
It doesn't matter if we're not the ones reading it.
If the LLMs are fine reading it, do we care so much about variable naming and all that kind of thing?
I think those are questions to be answered by understanding again that those higher level things, why should those things matter?
Well, it's things like, so why do we care about TDD and good quality code?
Well, we think it helps make the code more maintainable.
so that you can continue to deliver after you get that first release out the door, you can continue to deliver, make changes to it.
And when something goes wrong, you can diagnose it quickly.
You can find the fix and fix it without having unintended consequences.
Like we think those are the things that you will notice, that users will notice and that kind of business stakeholders will notice.
And so I think there's a whole kind of interesting bit of research to do over time to find out how much of those things, those lower level things actually impact.
on those higher level things that we care about.
So we may find that there are some things we still do care about.
Maybe we still do care about TDD.
Maybe it still does drive some good things that make the code maintainable, even when it's LMS maintaining it.
So I think it's more, I guess I'm giving you more of a question than an answer can in terms of what stays the same.
I think it is more that emphasis on, we have to get better at...
thinking about those higher level things.
Because we've always kind of, again, we've kind of been a bit lazy about it or just always struggled with how do we articulate to management?
Why should they invest in things to make kind of delivery better?
You know, things that we think make delivery better.
We've always struggled to articulate those things.
Now we have to articulate it to the LLMs and tell them to do it and measure whether they're doing it.
It's really interesting.
Great.
So I want to thank you both.
And I guess one little last tidbit for the listeners.
There was actually a section at the conference, and I don't remember the title of it, but it was something in regards to what's on the agent's reading list.
Do books still have a place?
And I can tell you that as a podcast host, there's only so much you can learn on a 40-minute podcast or on a 30-second YouTube short.
And so, you know, depth does matter.
And so please do check out the books by both Keith Moore and Andrew Harmel Law.
They are important foundational information and nothing's more important than a foundation when your building is shaking.
So again, thank you, Keith.
Thank you, Andrew.
It was great seeing you in person and I look forward to our next discussion.
Yeah.
Thanks again.
Thanks for having us.
Thanks again.
