# Harness Engineering: The Future of AI Code Generation

**Podcast:** Thoughtworks Technology Podcast
**Published:** 2026-08-20

## Transcript

Hello, everybody.
Welcome to another edition of the ThoughtWorks Technology Podcast.
It's kind of an interesting thing.
At ThoughtWorks, we have this thing called the Global Tech Leadership Forum.
And it's some of our technology leaders from around the world.
And we get together on a regular basis and chat about things.
And on a recent one of those, one of our guests put in the typed-in chat, you know, do we even need to see the code anymore if we're talking about AI?
which, of course, lots of people had responses to.
And so I'm sitting there going, you know what, this would make a good podcast because the podcasts are supposed to be discussions.
So I'm going to ask the guests to introduce themselves real quick.
Just going in the order on my screen.
Razin, can you introduce yourself, please?
Yeah, this is Razin.
I am a principal consultant in ThoughtWorks, which is also...
called as Enterprise Architect in some companies.
And I've been in ThoughtWorks for approximately 9.5 years now, almost 10 years, completing very soon, and working as architect in multiple projects.
And Kerr.
Hello, I'm Kerr Sanders.
I'm also a principal at ThoughtWorks.
My background is in a mixture of robotics and artificial intelligence.
I'm also a writer, so my degree in education is in writing of all things, which seems relevant today.
Very much so.
In fact, I may have to sign you up for a couple of blogs.
So it's interesting because, you know, this discussion came up and like everybody, I do initial research using the different LLM tools and so forth.
And so I went in and we have one that can see our corporate things, of course, and public.
And I said, hey, write me a little brief.
And it got the guest's roles wrong because I didn't give it the original chat.
It looked at the writings and the public thing, and it assumed that Care was the one that didn't care about the code.
And Razine was the one, from the enterprise architect perspective, that was going to be, oh, no, I need to approve everything.
I just found that really interesting that it came back.
So I guess I'll start with, you know, Razine, you started it with the question.
What's your stance?
Today, are you thinking today we don't look at code?
Is that a future thing?
Where do you stand on this?
What I feel is in 2028-ish, the code will not be any more relevant.
Why I think so is because what will happen is the tools that we are telling ourselves, like IDEs, etc., those will convert it into the new Harness engineering tools.
which would come up with lots of skills and harness and tools around it which will make spec so good to be converted to binary directly or if it is converting to code and then to binary i wouldn't be worried about the code so i would really not be worried about if the code is coming in between or not so i would just be worried about what specs i'd be want and the code is getting generated right or not and i will just generate the right test test it and just figure it out.
So you're not saying today, you are saying that's technically 16 months out.
Yeah, why I'm not saying today, even today it is relevant because I use lots of non-looking code things.
But the biggest problem today is token dynamics, etc.
But even today, right, for example, today, Meta has released one of the model called Moose.
And that is giving really, really good results in terms of coding.
So that started happening today.
And then methods like BMOD have started coming today.
But the biggest problem today is a token usage.
So that's why I'm saying not today.
When we will solve a token problem and local LLM problem at that point of time, and that is not very far.
This is going to be max one year or one half year.
And Kare, you said some opposing effects.
I think you were in a discussion that day.
with one of the leaders in observability.
I'm not going to use her name, but you're certainly welcome to do so because I wasn't part of that conversation.
What's your stance on this?
I've been working with AI for about a decade now and building an AI platform since before GPT-3 launched and brought AI upon all of us.
And so I've had a lot of experience with it.
And what I'm noticing is that AI-generated code gets the job done in the vast majority of the cases.
And I can't...
at all can test that.
But when I'm working on a safety critical system or a system where downtime costs millions of dollars an hour, I still have to read the code and make sure that I understand it and know what it's doing.
And I feel very strongly about that.
And so a lot of the work I do professionally is safety critical.
And so a lot of the work I do professionally means I have to read the code still.
But I was talking with one of these observability leaders in the industry and we've been writing a lot about sort of the split between AI skeptics and AI optimists.
And it's a very divisive split, like you're either one or the other.
And I find myself in the middle because the skeptics are like, we shouldn't use AI at all.
The optimists are like, we should use it for everything.
But what I'm realizing is that this transformation feels a lot like what we saw with infrastructure as code.
It's like 15, 10, 15 years ago when Terraform came out, there was this whole...
you know, movement around servers are no longer precious pets, their disposable cattle.
And that same analogy is now being applied to code where they're saying LLMs have made code cattle instead of pets.
But I think that's the wrong framing.
I think infrastructure as code made servers less precious.
And this AI generated code movement is making code less precious, which is different.
We still care about the code, but it's less precious as an artifact.
It matters more.
is all of the knowledge that goes into specifying it and i can go on along talk about that in a moment i'll get back to you yeah i have a counter to that right in infrastructure as a code infrastructure as code is a means to achieve a real servers which is an infrastructure but your real infrastructure is still a code whereas code is a means to achieve a real software that you want to achieve so what i'm saying is real software will really be important, still be important.
You will really need a software.
You will really need a test.
But there is one intermediary in between, which is nowadays coming, which the intermediary is called as code.
Basically, somebody is giving a requirement.
that is getting converted into epics or stories, etc.
And then one intermediary is coming, which is converting that to the code.
And then that code is getting converted to the real software.
And then real software is something that is being used in the world.
So what I'm saying is that intermediary's relevance will just go away.
It will be the first thing people will come write the epics, stories, specs, or whatever you want to call it.
There are different names that people are calling.
But that and then the final software and the test, those are the three things you will be worried about.
That intermediary thing that is code, which will be of no value and no meaning to the people.
Interesting.
So I try to not put my own opinion in here, but rarely successful.
In the time period that Kerry just mentioned, when they were talking about especially the code part and the other throwaway in cattle and things, some people may know I was deep in the DevOps movement and I was actually the product manager for our continuous delivery tool for many years.
And it was very easy, Rogine, to take that position and say, oh, you know, it's not really the same.
But we find that infrastructure is not just.
the hardware, right?
It's the networking layer and the performance and, you know, which data center and that kind of thing.
And often what we would do is we would put that on the development teams and say, hey, you're now a Terraform person, but don't worry because, you know, all the things.
And then suddenly we go down at 2 a.m.
and nobody would have any idea how to fix it.
Do you don't see that as a problem in that time?
Do you think it's just an economic problem today?
So the way we are right now caring in the Terraform about the real cloud, right?
Like, for example, AWS has its own language, which is a competitor of Terraform, and Azure also has its own, and GCP also has its own.
But we don't really...
Lots of people don't use those.
People use Terraform.
But while writing Terraform, they still think about nuances of which kind of a server I want or what kind of a thing I want.
So what I'm saying is while writing spec, people will be worried about what kind of a logic I want or what.
The logic that we write in code, the conditions, the loops we write in the code, those will be in a much better or a much simpler language that anybody can understand.
But people will still be worried about those in the specs that this is the logic I want or this is the category I want or this is the kind of observability I want or this is the kind of harness I want.
Those kind of things you will still keep writing on the spec.
But the spec and which we even today write in the stories and epics.
But stories and epics get converted to the code and then I'm just looking into those code.
Those worry will go away.
So all those things will be there in the spec.
And then you worry about the final software and you worry about the test.
Those three things you worry about and then everything else, harness will take care of.
So you're really putting all the weight in the harness?
Lots of, lots of weight on the harness and that is going to be a really big industry for sure because right now lots of harness is provided by...
the model provider itself.
But I see a future where model provider and harness provider will be two different industries.
And there'll be a really big weight on the harness provider.
And I see ThoughtWorks being one of the leaders in that industry, AIworks being one of the best harnesses we are going to have.
So we could spend an entire episode defining harness.
But luckily we did.
And it just published recently.
And so for our listeners, I will answer, we'll put links in the transcript.
There's a great article on martinfowler.com that talks about what a harness is.
Care what you're responsible.
What do you think about that?
Is a good harness good enough or guardrails good enough?
It's complicated.
When you're working with a system, so much of the system's specification isn't written down.
It's written and embedded in code.
It's written and embedded in people's heads.
It's all tacit knowledge.
It's not captured anywhere.
And so with harness engineering, what you're asking engineering teams to do is take all of that tacit knowledge and put it to paper in a very thorough way.
And so if you can do that, I do believe that that could work.
So the lady I've been talking to in observability charity majors, She recommended to me that I look at the Phoenix Architecture by Chad Fowler, which is this whole series that he's written on regenerative software.
And that series was the first time that someone convinced me that maybe it's a good idea to not let engineers write code anymore and only let them regenerate software from specifications.
There's a lot.
That could be a tone episode right there on why that is.
But if you truly can get to a point where all that knowledge is captured in specification, in documentation, in conformance tests that aren't specific to your tool stack, then I think you could generate all the code and never look at it and also nuke all the code every day and regenerate it from scratch if you had the budget for that.
Someone that our listeners may or may not be familiar with, I think all of the guests are, is Bridget Kromhout, who does a lot of DevOps things.
Bridget used to say, Kubernetes is easy.
People are hard.
Is that what I'm hearing from you, Kara?
It's like to get people to, you know, if you have a great spec, but you ain't going to get a great spec.
That's exactly it.
I think people and social problems are the stickiest part of engineering now.
And they always kind of have been.
It was just hidden behind a layer of writing lots of code.
Yeah, but there is one solution to that also and that's where again my harness engineering point comes into the picture.
I will divide spec into the two different parts.
Definitely people are not really going to really good at writing all those things into the spec or going into the details or as much as possible.
But what if we divide into the two different parts where we let some experts like for example ThoughtWorks write.
specs or nowadays it is called a skills and then people have also started calling it as plugins but we write a plugin which is a combination of skills and the skills are actually giving all the prompts or the specs about how you should write the code what are the good practices what are the things we should avoid what are the things you should not avoid all that is bundled into the skills given by harness providers and then the real creator of a software just gives what they really need.
It's all about what they need.
I mean, it's not, and that's my whole argument that why I think that the code is not going to be really important because in code, we do lots of things together, both the things together, good code versus what a really creator of the code means.
But here, the spec will only be what creator of the code means.
And then if they are not clear enough that what they need, So they will not get a clear enough software, but then they will realize, oh, I'm not getting clear enough software.
This is why my problems are.
So they will keep improving the spec.
But at the end, spec is what the real creator of the software needs.
Software engineering parts are handled in the skills.
Creator only says, this is what I need.
And they get that software.
That's going to be the future.
I need this.
I got that software.
If there is a problem on that, I got something else I did not explain very well, then I improve the spec.
And finally, I give the spec that, oh, I wanted the software.
I got the software.
These are my tests.
Done.
I just get the software and start testing.
Why would I care that what happens to that spec, how the harness provider is taking skills, how they're creating to something called a code, how they're dividing the code, good code, bad code?
Why would I worry about it?
So what about things like performance and carbon impact and all of that sort of thing?
If a non-deterministic system is writing the code and it passes all my tests, and, you know, most people don't know what fitness functions are as far as like response times and those sorts of things.
You know, what about that situation where the code that's created is just bad?
I mean, I'm, you know, taking a stance here.
I mean, I'm not saying it always is, but it's just not as performant.
And so now you're spending more time on your cloud costs and that sort of thing.
Are you worried about that or do you think the LLMs are going to be just better at it?
Why LLM have to be better at it is because if it is not good and then if people start seeing that, oh, this LLM is not providing good software, I'm just giving a prompt and then I'm getting a software, but it's not good enough, they will switch the LLM provider.
And the competition will make sure that people will start.
going to that standard that okay people really get what they want real performance code and then second thing is uh either it's it's a com it's not only the llm provider it's a combination of both llm provider and harness provider together we'll make sure that the code being created uh is really good and then people have also gone into the level of uh it's not only about people not caring code not caring about code it's also about i think there will be llm models which will directly create binary from specs Actually, LLM themselves also will not even generate the code.
Elon Musk said that I want to work on that direction, that I will directly generate a binary from a spec.
So that is a different direction and that will take five, six years.
But like I'm saying, the harness provider and LLM provider, if they're not good enough, then people will switch them and the competition will make sure that they're good enough.
I think that we may eventually get to the point where spec to binary is a reality.
I think it'll be a very long time horizon because there's so many architectures to compile to, so many architectures to target, and so many optimizations that lower level languages like C, C++, Rust are already doing today.
It'll be really hard to recreate, add an L, an inference layer.
That's one piece.
But the second piece is like a lot of my career has been spent in...
sort of performance engineering, I guess you could call it.
And, you know, trying to get things like vector search down to five milliseconds per request.
And so you start to really find what can you squeeze in your system to get there.
And I think in my recent experience, I'm using Claude Fable almost exclusively, which is a little wasteful maybe.
But I've been using it to build some systems in Rust on the weekends and trying to apply the sort of regenerative software principles to it.
What I've noticed is that even Fable, which is one of the most heralded frontier models, still does unnecessary memory copies and things like that.
And I mentioned that because those are the kinds of things that are really easy to not do.
And they tank your performance really quickly.
And so it doesn't mean that LLMs can't do that.
It just means that right now I'm still very much in a mind of I have to read the LLMs code.
I'm doing performance sensitive work because they're still making.
really junior mistakes in their engineering.
And that could get better over time.
It just, I'm surprised it hasn't gotten better already, in a sense.
I think this kind of problems will more be solved by harness side, the skill side.
Because if we are doing a performance related work, the skill needs to be keep enhancing that, hey, don't do this, don't copy memory, all those things.
Harness providers will keep into the skill and big skills will...
getting generated by harness providers.
So, and that's why I was saying, right, there'll be two different industries in future, LLM provider and harness provider, because both are equally important.
Writing into the skills that don't do this kind of mistakes and do this kind of coding.
Like, for example, ThoughtWorks saying that, hey, these are our practices, do these practices, do internally, this is how we generate code, small code.
All those will be our future.
We originally had another guest that we were going to have at the same time.
who's our regional CTO in Europe, Giles.
And one of the reasons we were going to have him is that he did some real research on the tokenomics or the economics part of this.
And he's going to publish all these results, so don't quote me on exactly.
But basically, he built roughly 150,000 line application.
A lot of it was refactoring from enterprise because that's a lot of what we do is enterprise modernization.
And then when he ran it, he noticed a thing go by in his terminal that said an edit to line 4000.
Scroll past his terminal.
He's like, whoa, what is this thing doing?
And so he took it out and he did some manual refactor to do it in stages.
and took his token costs from $159,000 to $27,000, 83% savings, because the tools he was using today just couldn't break it up into logic.
They could break it up into functions and that, but not logic, right?
They still act like they're reasoning, but they're not really reasoning.
And, Rosie, you said a second ago that it's the economics part that has to be solved.
What makes you think they're going to do that?
Why is that in their best interest?
Yesterday, I was working on one of the tasks.
I'll give you an example.
I asked LLM to do something.
I gave it and then it took 11 hours for a local model to do because it's very slow.
And it took 190 million tokens.
And then cost was zero because everything happened locally.
So we are seeing the examples of local model like QN 3.6 yesterday was able to do that.
And today Muse has been released, which is claiming to be even better than what Quen 3.6 is able to do that.
And this is what companies need.
Companies are not going to spend millions of money on API costs.
So they have to be reduced.
Otherwise, nobody will use cloud, right?
If it is not as usable as everybody wants it to be using, then they will go bankrupt.
They have to reduce the cost.
and then many companies can't afford so much api cost so people will go into the hardware side people will go into the local model side so that's why i'm seeing the trends i'm not just guessing it i'm seeing the trends that good models are coming and people are moving to those direction and if i'm able to do like a like a full learning management system kind of a coding in 48 gb ram then i'm pretty sure when 128 gb becomes common for everybody or or terminals become common for everybody, I think we'll be able to do much better.
Is this a case though of the rich get richer?
So what do I mean by that?
Not everybody can afford to run things locally.
Not everybody can afford to run 128, you know.
I have a friend doing a desktop machine and with the cost of RAM, I think it was $450 US dollars for 32 gigabytes of good quality, high performance RAM.
And I look at another experiment that was run by someone in my office at ThoughtWorks.
The tools like using the correct model in like Cloud and OpenAI's models and those sorts was $5 to $7 to create this application that she wanted.
There was a low-code, no-code platform that's prompt-driven that I don't want to name, of course.
And it was $27 to do it.
And so the cost to the person who can't afford the local, it's kind of like the same thing happens with food, right?
If you're in a neighborhood that's a food desert and you can only buy groceries at the local corner store, you pay $6 a gallon for milk.
If you're in an affluent area that has the fancy store, you pay $3 a gallon for milk.
And so how do we as technologists make sure that societally, everybody can do this, right?
Or is it going to be only the people that can afford local models?
Or how do we get it to be the business people who, let's be honest, they're creating software and they're creating bad software in many cases.
How do we get them to do this?
Do you think it's a tool problem?
It's a hardness problem.
The hardness is not matured enough to have skills that are giving LLM enough ways of doing the coding.
Like, for example, BMAD method, right?
I don't know the creator, so it's not a promotion or something.
But BMAD method is something which is right now a combination of skills where it is doing the full agile.
You give a prompt and it is creating a simple...
PRD, then UX document, then architecture document, and then it is creating what do you call the code and then review code.
And then they are talking each other, two million columns are talking each other and solving the problem.
So that kind of a harness when it will come and more people will look for reducing the cost.
If the cost of the tokens will reduce or local hardware will come, then that kind of a competition, right, will be making things affordable for everybody.
So better harness will reduce token cost and competition will reduce token cost.
And then the competition between local model and competition, local model and API when it comes, right?
Automatically, entropic or open AI clients will reduce the cost of the tokens.
Kara, what are your thoughts?
I am always cautious about playing around what a business team is going to do to lower my costs when their incentive is to make as much money for me as they can.
So I'm less optimistic about AI costs going down.
If anything, I couldn't go up as all the subsidies and incentives start to dwindle.
That said, I've noticed that I have peers who have built software using Claude.
Let's say it was Sonnet 5 or something.
And they've spent genuinely thousands of dollars in API calls a month to build a relatively simple application.
And I'm able to do the same thing with a frontier model for like $10.
But the difference in our workflow is that one is very waterfall where my peers are saying, here's my entire spec.
I'm going to write it ahead of time.
Go one shot it.
And the other one is I'm writing a bunch of tiny concept documents.
If here's all the major criteria that I have for the system, the sort of externally zero behavior I want it to have, let's do one at a time slowly.
That takes a little bit longer, but I get a better result for a lot less money.
And it's like, that's my own sort of compromise when thinking about LLM cost is just changing the way I use it as opposed to hoping for external change.
So one of the things that I've poked a little fun at some of the LLMs on LinkedIn and other places, like I had asked a question eight or nine months ago just for correlation between payroll and success in U.S.
baseball.
And I said, hey, take the World Series had just ended.
Take the top five teams.
Show me what their payroll was.
I just wanted to see if there was a correlation.
And I knew that it might not be accurate on exactly the payrolls because, frankly, they make that stuff hard to track.
For those that don't know how U.S.
sports work, they give them free cars and houses so they don't have to show it up as payroll.
So whatever.
But what struck me is it got four out of the five wrong on where they finished.
I mean, this was easy to find information of who won the, we call it the World Series, even though it's only the United States, because that's just the way we are.
And the other, you know, top four on top of that.
And it was, it got one of them placed right.
The other four, it was wrong and where they even finished.
And so, you know, I think about like if we watch the media and if our listeners, if you watch the media, if they cover a story that is something that you know really well, you're going to be like, ooh, that wasn't really right.
A lot of the time.
If you do a query to most LLMs about something you know really well, a lot of times you're like, ooh, I have to tweak that.
You know, like I'm going to ask Claude to do a function, a small piece of code, 20 lines or whatever in your IDE.
And you're like, ooh, no, I'm going to change that.
So if we are experiencing that level of accuracy, that doesn't worry you, regime?
Are you just going to be like, as long as the test pass, I'm okay?
There is a concept called bug in the software development, right?
What is bug?
Bug is something that I wanted some software.
and I did not get that software.
So same thing is going to happen and how we handle bugs, same way we are going to handle in the future.
So if I'm not going to get the accuracy or the things that I wanted, I improve my spec that, hey, I wanted this also, which I am not getting.
And then that, consider it as a bug fix.
That will go to the spec.
Spec will get improved and then internally it will do.
like code or whatever it wants to do, but it will fix the bug.
And then it will come back with that accuracy.
And if it doesn't come back, then it's LLMS problem.
LLMS will solve it because I explicitly ask it to solve it, that I need this accuracy.
So this is how it will happen.
But yes, on the first go, even today, we don't get a software that we want.
We get things like bugs.
We get things like improvements.
We get things like MVP.
Even in future, we will never get it.
Most of the time, even humans who really want the software to be created, they themselves don't know what they want to be created.
So they will give halfback requirement and they will get halfback software.
They will keep improving.
They will keep improving.
Yeah, I have to admit, I was playing devil's advocate there because we recently, we were talking to one of the national health services for our country.
saying, hey, this AI for medical was only right on a initial diagnosis like 85% of the time.
And, you know, 15%, that's a lot.
It's wrong.
And it was funny because the CTO stopped and looked at our CTO and said, well, how accurate do you think people are?
It was 82.
And so it turns out medical diagnosis is hard.
Is there a tendency, I guess, and Kara, I'm going to ask you this because you were talking about people overthink.
Is there a tendency that you've noticed where we expect more out of these tools?
We're like, oh, this wasn't perfect.
The person maybe isn't either.
The phrase I'm thinking of is cognitive surrender.
I don't know if people have higher expectations of LLMs than humans.
That's a little unclear to me.
Because, you know, like this example, right?
We said 100% for the LLM, but people are only 82%.
That could be the case.
But what I have noticed is that people are more inclined to trust an LLM over a human, given all other factors being equal.
And this has been used maliciously.
You've had...
Leaders sort of come back to their reports and say, hey, Claude said this, you said that, I trust Claude, not you.
You have reports trying to trick Claude into saying their argument to their leaders so they agree with it.
Like all kinds of things going on around trust and trust being disproportionately allocated to the machine.
And I think that's, for me, the scary part about LLMs in general is that we just assume that because it's a system, it's deterministic, it's trustworthy, it's audited.
But LLMs are kind of the antithesis of that entire idea.
So I would definitely agree.
But that said, does that apply to code creation?
That certainly applies to give me advice about my family or give me economics that understand my point.
But does that expand to our particular topic of code creation here?
So for code creation, what I've noticed is that I have a personal bias to trusting an engineer more than I trust an LLM when it comes to generating code.
We had a funny incident a few days ago where I was reviewing a code base with a team that contains a lot of contributions from AI, from people, all kinds.
And I was looking at a section of code that we all thought was pretty good.
And then I was like, this section of handwritten code.
And we were all like, how did the model know it was handwritten?
We never told it.
What got it there?
But say that aside, we learned there's a subtle difference between what code that was written by a person, code that was written by the machine.
And I trusted a little bit more.
Now, do I trust code written by a junior developer who's in their zeroth year over Claude Fable?
I might actually, if they were writing it by hand, because I know that they knew my intent and were trying to follow it to the best of their ability and not hallucinate it along the way.
Whereas even the best model can hallucinate and get stuck in a rabbit hole.
Yeah, I wouldn't even worry about the code.
I mean, if it is right.
I mean, it's all about do you trust more to human writing the code or cloud to writing the code?
And I am saying that it doesn't matter who writes the code.
I don't even care about the code.
What I care about is I wrote a spec.
This is what I need.
And I got that software by testing it.
If it is working what I want, I don't care.
Good code, bad code, what you wrote, what you don't wrote.
And then you shouldn't be charging me for more changes.
I mean, that's a harness problem.
Otherwise, I'll switch the harness provider.
So I wouldn't even be worried about something called a school.
So it's interesting because you've used the phrase harness provider many times.
And most of the time I've run into the term, now I don't do code day to day like both of you do.
Okay.
But most of the time that I've run into the term, harness is something that that team is creating.
There's not a harness provider.
This term is being misused a lot.
So when we are creating agent-related work beyond coding, then we are the harness provider where we are using agents, etc.
But this term is also nowadays being used for IDEs, something that we used to call IDEs, like VS Code or Codex or Cursor, right?
Why they are called as harness providers nowadays?
Because they are doing exact same thing that we are doing when we are creating a harness for agents.
So nowadays they are doing it.
and they are doing it for the agents who are doing coding.
So that's why IDEs have started being called as harness providers because they provide harness on top of LLMs on how you should do coding, what you should do, what you should not do, what practices you should follow.
All those skills, plugins, etc.
are provided by companies like Cursor or Codex.
So that's why they are called as harness provider nowadays.
So like if part of the harness is their security rules, those would be a skill that you would get from Claude and not written by you?
Yes.
And then we are also getting some Python scripts, ready-made Python scripts.
There's some ready-made skills.
Those things are coming nowadays with those tools like Codex.
Because Kara did use her name earlier, I will quote Charity Majors.
a talk that she gave at a conference I hosted and say, that scares the hell out of me.
I'll tell you one harness provider name.
You will be surprised how this word is misused.
So one harness provider is called as AI Works and the company name is ThoughtWorks.
I think I've heard of them.
So that's going to be a thing.
I mean, people will give their skills.
their plugin, their scripts into something called as harness, and then you can use any other LLM on top of it.
And then you code on.
So there'll be three parties here.
One is a harness provider, one is a spec writer, and then one will be the actual LLM provider.
So three parties will work together.
So one of our colleagues in the discussion, in the thread, who's, again, very much on the infrastructure as code side, had said that also doesn't really care about the code, but was mentioning things that you don't typically see in a harness, like a fitness function.
So for those that aren't aware, the very short definition of a fitness function is a test that tests for a business requirement, like must respond within this many milliseconds, must be able to run on this GPU, or what have you.
And so it's something you can put in your CI pipeline to say, No, you're not performant enough, even though you're functionally doing what you want.
Does that mean that things like fitness functions, whether you call them that or not, need to be part of the harness now?
It will be provided by harness and it is part of AIworks.
So the harness provider will provide that for you to create a software.
Okay.
So again, that's not in cloud providers' best interest financially.
That's why I'm saying those two will be two different industries.
Kher, what do you think?
Do you trust them?
You know, harness is one of those phrases that it's a term that's capturing so many different concepts and bringing them together that it almost feels too big to me.
Like, I think there's different kinds of testing and compliance that we're looking for with harnesses.
Like, some of it is guiding the LLM along its path to generating the right thing.
The other half of it is once the thing has been generated, being able to validate that thing and determine if it's the right.
version of that thing.
It's two different ends of the problem.
And so for me, I think you mentioned fitness functions, there's property-based testing is becoming a big thing.
For me, I do think some iteration of harnesses are going to be the future for LLMs, but I also think part of it is the way the models are trained themselves.
Because really everything we're doing is trying to sort of one shot or two shot on top of a prompt engine, as it were.
And what would be even better is that the model was just trained from the very beginning to write code that was good in a certain discipline or a certain industry instead of having to give it tons of prompts on top and kind of waste your context window.
And LLM caching can help there.
It can.
It can.
So you mentioned, Razine, that you think this is 2028.
You already mentioned economics.
Kerry, you've mentioned some things around human dynamics and so forth.
So I guess I'll start with you, Kerr.
What has to be true for you to trust this?
What would have to change and be truth on the thing you're working on so that you would trust it?
One very simple test for me would be if I can go several months or even a year working with a model and not look at it and be like, wow, you really hallucinated that requirement.
But it still happens regularly enough that every time it happens, it erodes my trust and makes me more eagle-eyed and looking at it again harder and harder each time that happens.
So part of it is just, can I go a span of time without being let down?
And the other half of it is, can I get myself, I've been trying really hard to experiment with writing code without looking at the code too much, which is very difficult for me.
And what makes me more comfortable with that is, can I get to a point where my, conformance testing, my property testing, my specifications are so robust that I can delete my repository and recreate it from just that harness and get the same result at the end.
If I can get to that point for myself, I can at least hit myself control with the idea of trusting AI to generate code in the context of my repository, just not in the general context.
Define for me, if you will, same result.
Like if it completely switched languages, but met functionally the spec?
What is the same result to you?
Yeah, so same result as the externally observable behavior.
So like I used to write patents for distributed systems.
And a lot of my work was spent taking this patent spec and implementing it in a language.
And I went from like Python to Java to C to C++ along the way.
And all of them were the same system, just different attempts at getting to fulfilling that system in the spec.
And so I see...
this workflow with AI very much the same way where I can write a very detailed specification for a system and the AI can implement however it wants as long as the external behavior meets the spec.
And Rosin, anything besides economics?
Are you there already?
Is it already, you know, what has to be true?
Or are there other technical things that have to be true?
I don't think that we have so many matured skills yet which will tell LLM to do the right coding.
Maybe LLM might be really good but skills are still not there.
Like for example, why ThoughtWorks is famous for doing some sensible defaults and some things.
But there are no skills available which are saying that this is the sensible defaults that LLM should follow on the market or in the industry.
So once those skills will be available that this is how you should do for LLMs.
then somebody who wants to write the code, they will take that skills and then use it into the LLM with their intent of writing the code.
And then they will have a better code compared to what they are having today.
And then you already mentioned token economics.
Those two things should be really changed.
So those harness providers are not there yet.
Those skills are not there yet.
Sorry, I will not use harness provider now.
But those skills are not there yet.
Those ready-made skills are not there yet.
This is how you should write a good code.
This is how you should test the code.
This is how you should worry about the code division, the files division, the reusability.
Those things, right, are still not there.
Great.
So I want to say thank you to both of my guests today.
It's been a really good conversation.
I have to admit, I'm looking at the clock saying, oh, I have to stop.
We've got 30 more minutes.
Easy.
You know, maybe even more.
But maybe we'll do a follow-up episode.
That said, just another note to the listeners that these are complex topics and we like to do podcasts to talk about, you know, what are the things that get you thinking about it.
We know it's not deep enough for you to go in Monday morning and change your behaviors.
But there are several like martinfowler.com that will link talking about harness engineering.
I believe Razine is going to write a blog on this that may be live by the time this podcast goes live.
And so I encourage you to continue to research this area.
It's not a solved problem, I think, is what I'm hearing from everybody.
But Karat Rajin, thank you so much for your time.
I really appreciate it.
Thank you.
Thank you.
Thank you for inviting us.
Thanks a lot.
