# AI Code Generation Requires Industrial-Grade Evaluation Harnesses

**Podcast:** Software Architektur im Stream
**Published:** 2026-09-08

## Transcript

A note from our team.
Software Architecture im Stream will be streaming live from the ISAQB Software Architecture Gathering in Berlin this November.
Join us on site.
You'll find more about the program and special discount code for our community on our website, software-architecture.tv.
Hello and welcome to another episode of Software Architecture im Stream.
Today with me is Randy Schu.
Randy, may I ask you to introduce yourself?
I hope I pronounced your name the right way.
Thank you.
Actually, it's a German name.
And so as we were chatting personally, it's Schaup.
Even though it's spelled incorrectly, we can have that conversation later.
Rolf, it's great to be with you.
So yeah, hi, I'm Randy Schaup.
Right now, I'm the head of engineering for CircleCI.
Wir sind ein Kontinuos Integration Service Provider, wie viele von Ihnen wissen.
Ich habe mich über die Software gearbeitet, über 38 Jahre, also eine lange Zeit.
In der meisten Zeit, mehr als halte ich mich als Architekt in einer Form oder anderen.
Ich habe mich als Architekt an Ebay, als ich als Architekt an Ebay bin, als ein Security Company.
Und wenn mein Titel hat Architekt in es oder nicht, ich habe immer noch liebte Software Architecture, Distributed Systems und so weiter.
So, again, ich habe bei Ebay, ich habe bei Google und ich habe bei einem anderen Silicon Valley Unternehmen gearbeitet.
So, ich hoffe, wir sprechen über all diese Dinge.
So, was ich nicht gesagt habe, aber ich habe herausgefunden, dass, ja, als ein kleiner Kind, You already had contact to the newest technology at Xerox Paolo Alto Research Center, Xerox Park.
So you were raised with technology, the newest technology.
So, yeah, it makes sense that you worked at those huge companies.
So how was your time when you played around?
I know we didn't talk about this, but...
Ja, ja.
Ich wollte es, zu haben, die neue Technologie zu haben.
Wissen Sie das war wirklich etwas, was andere nicht mehr access zu hatten?
Ich habe nicht.
So, die Geschichte ist, mein Vater hat sein Ph.D.
in Computer Science in 1970.
Und er hat es...
der Carnegie Mellon University.
Das ist ein sehr guter Computer Science University.
Und das war einer der ersten Ph.D.
in Computer Science in den USA.
So, obviously, die Leute haben sich in Computer Science für Jahre alt, aber das war die erste Zeit, das war eine sehr viel Studie.
Und er und meine Familie kam aus dem San Francisco Bay Area, die sich in Silicon Valley later auf angefangen hat, um das Werk zu machen.
Mein Vater hat der Xerox-Research Lab in 1971, so nicht sofort nach dem PhD, aber ein Jahr später.
Er war in der ersten Gruppe der Arbeiter.
So andere Leute, die bei der gleichen Zeit waren, waren Alan Kay, der in der Smalltalk, von course, Butler Lampson, Chuck Geschke, der in der Adobe, Let's see, the guy who founded Microsoft Word, whose name is escaping me at the moment, Simone, Charles Simone.
Anyway, a whole bunch of great people.
Yeah, it was, and the serious computer science in the early 1970s was very small.
Everybody knew one another.
I can tell you a story about meeting Don Knut, the very famous professor who my dad and his friends like.
Ich habe mich sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr, sehr Es war weird für mich und mein Bruder zu begren, mein Vater zu gehen, um die Wochen zu arbeiten, die wir absoluten haben.
Weil wir sehr klein waren, wir in die Beanbag Conference Room waren.
So ich hatte eine Conference Room, ohne Tafeln und Tafeln, aber nur Beanbags.
Und wir würden machen Forts und spielen, wie Kinder.
Aber dann mein Vater war in Computer Graphics.
Und er hat sich eine der ersten Paint-Programs gemacht.
So wenn du Mac Paint oder Microsoft Paint jetzt mal, All those metaphors were stuff that he and his colleagues invented in the early 1970s.
So having one part of it that's the canvas where you draw and another part which is the palette where you choose your color and your brush.
You choose your color and your brush and you go and you draw on the canvas.
And so only later do I realize that my dad's research at Zurok Park was the only research that would be interesting to a six-year-old.
Weil sie andere Dinge sind, die erste Wortprozessor, die LaserPrinter, die Ethernet, die Object-Oriented Programme, die Touchscreen, die VLSI Logic, die ist ein Weg zu legen, die Circuits auf eine Chip.
Ich habe ein paar Minuten, aber das Group an Xerox PARC in den letzten 70s hat, die in den letzten Jahren in den letzten Jahren, die in den letzten Jahren, Computer Science right at that time.
But it was lovely.
I really enjoyed drawing spaceships on my dad's paint program.
So you grew up in the middle of the heart of technology where everything was invented.
And I guess there were moments where you noticed that some technology now reached the consumer state where you said...
Hey, that was something I played around with some years ago.
Yeah.
Is it really that way?
It was, really.
So there's a much longer history, obviously, associated with Xerox PARC, but just very briefly, the goal of the lab was to build the office of the future and essentially build a personal computer.
Now, remember, we're talking 1970 to 1973 here.
The IBM PC didn't come out until nine years later, 1982.
The Mac didn't come out until 1984.
And it was really only the Mac in 1984 that had the same combination of things that they had developed at Xerox PARC.
Again, graphical user interface and WYSIWYG.
There you go.
WYSIWYG work, essentially.
What you see is what you get.
Und das war die Hauptsache der Grundlage.
Und wieder, die Laser Printer war Teil, Ethernet war Teil Teil, Object-Orient-Programm war Teil Teil, und die Graphical User Interface war Teil Teil.
Mein Vater war unerhöhnt, in Computer Graphics.
Aber die Alto, die sie in der Zerox PARC gebaut, wurde der Mac.
Und da ist eine wunderbare Geschichte über...
Steve Jobs, getting a tour of the lab in 1979, and by some accounts stealing the work, by other accounts being inspired by the work.
It doesn't really matter how you want to talk about it, but it is a direct line in our industry from the research in the Office of the Future and Personal Computing at Xerox PARC into the Mac and then Microsoft Windows and then how we use computers today.
Fascinating.
So you experienced all of this and then did your own way, worked for large IT companies.
And now we have the age of AI, if I might say now.
So for, I think, three or four years, we now experience all the LLMs, how good they are.
Many people said, hey, it's giving me the wrong answers.
I can do better development than I get those auto-corrections and so on from the AI.
And it evolved.
And so now how do you experience how the AI LLMs evolved?
What's your current take on it?
Ist es wirklich schon schon sehr hilfreich?
Ist es wirklich AI?
Oder ist es noch immer noch ein Stochastic Parrot?
Ja, Stochastic Parrot.
Das ist eine wunderbare Fähre.
Obviously, es ist nicht mein Fähre.
Ist es AI in der Sinne von ist es human-level intelligence?
No, nicht yet.
Ich habe viele Gedanken über AI.
Und ich finde, dass ich pragmatisch über das, als ich komplett pro oder komplett anti bin.
Ich denke, es gibt mehrere Dinge.
Nummer eins, AI als ein Tool ist absoluten, die unsere Industrie zu transformieren.
Es gibt keine Argumenten über das, eine Art oder andere.
Ich denke, es ist nicht ein Professioneller Software-Developer und komplett ignorieren AI.
You can use it more or less, but you can't ignore it.
It exists, just like electricity exists, just like compilers exist, just like a bunch of other tools that we have in our world.
It is very clear that the mechanism of typing in code is AI is able to do that.
There's no argument in my view that AI can take, again, through experience, AI kann man einfach nur einfach nur einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach einfach Wir sind die Industrie- und Software-Engineering, die wir nie hatten vorher hatten.
So, wie in der Industrial Revolution oder der Agricultural Revolution, die Entwicklung der Technologie hat so gemacht, dass wir Dinge manually machen, und dann haben wir die Maschinen, die wir machen, nicht all die Arbeit haben.
Ich denke, wir sind das als Software-Engineering.
So, die LLMs sind zu tun, die wir machen, die wir machen, manually, typing in code.
Does that mean that humans and architecture are not useful in software engineering?
It doesn't mean that at all.
At the same time as AI is doing all the typing work for us, at the same time what we need to engineer is the system to make that safe.
What do I mean by that?
AI, LLMs produce code at a rate no human can keep up with.
Was ist jetzt die engineering problem?
Das ist nicht die Typing.
AI kann die Typing.
Die engineering problem ist für uns das zu machen, und das ist die Ausstellung der Evaluation der AI-produzenden Code.
Das ist Static Analysis, das ist Linting, das ist Static Typing, so dass die Compiler kann checken, all die Worte zu haben, dass Adversarial LLMs ...
LLMA produces code, LLMB reviews it.
That is spec-driven development, where we say as a spec what we want the LLM to generate, and we check the spec via tests and via evals at the other end.
There's a lot more to that evals, but I'll just stop there.
I think AI is a good producer of things, and now the engineering problem is channeling and almost back pressure against the ALM generation.
But you just named all those things we learned during the last 20, 30 years, how to do good software engineering.
And somehow it seems it was lost.
Not everything was used in software development.
And now we notice that Hey, it might be a good idea to write a good proper spec and to have some tests and static code analysis to work with the LLM together.
Isn't it funny?
It's funny and also wonderful because LLMs are telling us in a new way all the things we knew about building good software.
It's forcing us to build software in the correct way.
So, make it clear that I'm not missing anything.
That is to your point, test-driven development, behavior-driven development, where we specify the behavior in an executable form.
That is writing a spec ahead of time and validating that at the other end.
That is small iterative forward movements instead of let's type for six months and then think we're done.
Produce incremental.
und check that incremental value or iterative step along the way.
Literally everything that we know, we, this architecture community, have known, maybe frustratingly, for decades about how to build good software, AI is reteaching us that.
Why?
Because we have...
How could we get away with doing software without these verification steps?
It's because we were implicitly doing it in the software engineer's brain.
Whether the software engineer was actually writing the tests or not or doing it in a formal way, the only way we got away with it was because we had a human in there and the human could manually, in some way, manually do the verifications that we're talking about.
And obviously the more we let the computers do it, the more we use tests, the more we use static code analysis, the more we used compiler stuff and metaprogramming, the easier and easier.
AI gives us the power of producing the typing without the human judgment.
And so we need to encode that judgment explicitly into the evaluation part.
Does that make sense?
Yeah, it makes sense.
Someone in the chat writes, he or she has always distinguished between the role of a developer and that of a programmer.
Wenn wir punch cards hatten, war es sogar einfacher zu distinguieren.
So AI ist jetzt der Programmierer und wir sind die Entwicklung, so wir versuchen, zu sehen, die große Bildung und denken über die Aufgabe und die Aufgabe ein Ziel.
Die Aufgabe, die AI, wir geben es ein Ziel.
Ist es etwas, wir können das so sagen?
Das aligns mit dem ich.
Ich glaube, completely.
Ich glaube, es ist ein, wenn wir es ein Programm oder ein Developer oder etwas sagen.
Ich sage das, ich glaube, mit dem sentiment.
Die words nicht wirklich matter.
Es ist jetzt frei oder nicht frei zu produzieren.
Das bedeutet, dass wir das nicht fertig sind?
Nein.
Denn das ist das Code auch nicht?
Does the code do what we think it does?
Does the code break everything else?
Does the code violate invariance?
Does the code violate the architecture?
So all these questions are now explicit things that we need to encode in an automated way.
Why in an automated way?
Because LLMs produce code at machine speed and we need to be able to evaluate that code at the same machine speed.
Number one.
Number two, I know everybody knows this, how did LLMs even learn?
It's feedback, right?
It's RLHF.
It's like, I'm forgetting my words, but like, because I'm tired, it's a learning loop.
The way that LLMs learn anything is through this loop and the way that they're going to learn how to do our stuff is a loop that we need to produce.
So as LLMs produce iterative new codes, new PRs or whatever, we need to have evaluations that are right there giving that LLM direct feedback, both to stop the LLM from doing things wrong and violating stuff, but also to teach the LLM to do it better the next time.
But LLM is already trained.
I can't really modify the weight.
So the teaching is in terms of setting up coding loops.
So I noticed that, yeah, it doesn't stick to my linting rules.
So I will give it a linter so that it can check its rules.
Yeah, that will always be true.
Die LLMs sind besser und besser, natürlich.
Und auch, keine Modell weiß genau unsere Code.
Und, übrigens, ich weiß nicht, dass wir uns über die Frage stellen, aber, wenn wir ein Schritt zurückgehen, ist, dass wir immer wieder ein neues System machen, wir machen etwas neues.
Ja, ja.
Das ist nicht genug.
As I was trying to say, we need to encode all of the concerns that we care about.
So again, static code analysis, compiler checking, tests for functional adversarial evaluation, security vulnerabilities, architecture.
I'm missing some, but everything that we care about that we would do in a human code review, we want to encode that.
Oh, and here's what you were saying.
A way to do that is to continue to try to teach the LLM up front not to make mistakes.
We should do that.
And also, that's never going to be enough.
So we always have to have an independent evaluation of what it does.
And we can do that independent evaluation with other LLMs, right?
So an adversarial LLM A produces some code, LLM B checks it.
We can and should do that.
But also, an even more...
Efficient way to do it is to every time we can make a deterministic, not a deterministic check, a static code analysis, a traditional code way of evaluating something that is both more deterministic, so more likely to be correct, and also way cheaper, like by orders of magnitude cheaper.
It's like, you know, LLMs.
wie ein Mensch, in einem senschen.
Der kostet, so wie ein paar Dinge wir nicht sagen, dass wir Menschen nicht tun haben.
Wir produzieren Chips in Fabrikation Faciliten.
Every single Chip produziert von Intel oder Samsung oder TSMC hat sich zu verabschiedet, aber nicht von einem Mensch.
Ja, aber wenn du über die Verabschiedung des Verabschieds, ich meine, um die Verabschiedung des Verabschieds, die natürlich auch zu den Harnessen können.
Sie können in eine sehr schnellste Weise beitragen.
Die LLM bereits hat einen Compiler, die Syntaxe checken.
Und andere Sachen gehören zu den Build Pipeline, weil sie länger dauern und nicht mit jeder Request wird.
So, es scheint, dass du jetzt gerade am Anfang dieser Request ist.
wie wir die Dinge behandeln.
Ja.
Und ich finde es faszinierend zu sehen, was die Partie in der Harness sollte, und wie wir das Ganze umgehen, so dass wir erkennen, wann wann, was kind of test zu gehen, um die Pipeline zu gehen.
Ja, ich kann...
Fantastic question.
Let me see if I can, I'm going to restate it slightly differently because it's part of the answer.
Our goal, our goal is ultimately to produce customer value.
And the way that we do that is producing good software.
How do we produce good software?
There are, as you were talking, I can think of at least four points.
I want to type it right the first time.
Yes.
Right.
So that is context engineering.
That is a good LLM model.
Das ist Kontext Engineering.
Das ist ein guter Speck.
Probierer andere Dinge.
Type es richtig.
Dann gibt es, wie es das Kompiliert wird.
So, es gibt Dinge, wie die Typen in die Bindungen, die bereits in die Bindungen sind.
Dann, nachdem, da ist, was CircleCI und andere Menschen in der Industrie sind, die Inner Loop.
Ich werde die inner-loop und die outer-loop sprechen.
So, die wir type was wir wanted der ersten Mal.
Die wir checken es via compilen in der exacten Moment, vielleicht linten als auch.
Und dann ist der andere Teil der inner-loop, was die localen Tests können und localen Evales können, sehr, sehr schnell, sehr schnell, so am I am der same speed, wie ich die Daten produzieren kann.
Und das ist die inner-loop.
All of those together.
What you point out correctly, and this has always been true, and it's still true, is there are things I want to check that are too expensive to check in that inner loop.
That's the outer loop.
And the outer loop, we can call that CI.
And like, don't do it for me.
But like, even if I didn't work at CircleCI, there always was and is today in software.
That outer loop where I've done as much...
As efficiently and economically as I can, I've checked my work as much as I can in this inner loop.
And then now there's more tests that take longer or take more resources or, you know, for various other reasons, need to run in a special environment that needs to run in that outer loop.
And you ask a very good question, which is where do the evaluations run?
And the answer is, I can give you the hand-wavy answer, which is.
As many as possible should run in the inner loop, as can be efficient and economic.
And then if it's not economic there, and it also matters, you do it in the outer loop.
And the reason why I give you that hand-wavy answer is because what is and isn't economic or efficient is actively changing.
See what I'm saying?
Yes, yes.
I feel what you're saying.
For me, the outer loop is a slow one, but...
Every developer has the same build loop, so build pipeline.
And that's a benefit.
And with the inner loop, the local build pipeline, I get a feeling like snowflakes.
Different developers do it in a different way.
So on a team, you have different quality of source code, which enters the build pipeline.
It might even be something that...
One developer tells the system, hey, remember not to make this mistake anymore.
And it lands in the agents MD of this one developer.
And so, yeah, you get those snowflakes.
Yeah.
Which are slightly different.
Yeah, yeah, I agree.
And so we're going to violently agree about this.
There is, to your point, as good as...
Good as we can make the inner loop, we still will always need the outer loop.
Right?
I agree with that.
We need the outer loop for several reasons.
To your point, my inner loop might be different from yours, and so maybe we don't trust that they're the same.
That's legitimate.
Also, to your earlier point, there are checks that we have to run that are not efficient enough, fast enough, efficient enough.
Es ist auch nicht so, dass es eine Art Erlöp ist.
Aber was wir sehen über Zeit, und diese sind alle kompatibel Ideen, was wir sehen und wir werden sehen, in meinem Erfahrung, ist wir werden nachdem wie viele Evaluations wie möglich.
Und ich bin wahrscheinlich in der wrongen Richtung von Ihnen, aber die Outer Loop, think of es als ein...
of a workflow.
And the earlier and earlier and earlier we can do an evaluation, the better.
That's for sure true, right?
Because we get better feedback, it's closer, it's easier to modify when it's earlier in this pipeline.
So that's why my answer, you know, you were like, when do we do which checks?
Do it as much as you possibly can in the inner loop to the left.
Und da werden immer mehr, weil du auch immer ein Artifact hast, an dieser Stelle, in der Outerloop.
Und das braucht noch ein paar Einstellungen auf den Artifact.
So, was ich verstehe, ist, dass die Outerloop ist also eine Artifact-Alignment-Layer.
So, dein Code wird die gleiche Artifact-Lintern beobachtert.
Und so, die Linter wird die gleichen Regeln für beide uns testen.
If, for instance, I don't use a linter and just let the code generate whatever, the LLM generate whatever code it likes, then my LLM will hit some boundaries in the outer loop, which it has to correct, will run more correction runs than your LLM, which already has a linter, or maybe...
Es war bereits trainiert in eine andere Art, so dass es nicht mehr braucht, es einfach nur eine Art, die Code aus dem Motivation zu schreiben.
Ja, und ich glaube, wir sind von viziently agree, und ich kann sehen in den Chat, dass die Leute auch agree, das ist großartig.
Ich habe jetzt noch ein paar Numbers zu diesem, das ich jetzt wissen, an CircleCI.
So CircleCI ist, again, wir sind ein CI-Provider.
Ich habe da für alle zwei oder drei Monate, aber...
CircleCI hat sich schon seit 14 Jahren oder so vorhin schon seit Jahren.
Wir machen jedes Jahr eine Studie, die die State der Software Delivery ist.
So die State der Software Delivery Report.
Ihr könnt es und downloadet jetzt.
Ich werde dir über das, weil es super interessant ist für diese Gespräche.
Weil CircleCI runn viele Leute'e build-pipelines, wir sehen Das ist es nicht die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, die schnellen, Es ist wirklich fast.
Und Ihr nächst Frage, ich hoffe, ist, was in der Welt sind die Unternehmen das Neues-Tempen-Kammerkne machte?
Und per die Leute in der Chat, es ist, die Bild-Pipeline so dass du die Opportunität und die Impacte von Merge-Conflicts verhindern.
Was wir sehen in den Average ist, ich bin forgetting die Numbers, aber es ist in der Report.
Ich mache das Nummer.
Vielleicht ist es zwei oder drei oder vier Tries für ein Change zu passen CI und werden in den Produkten oder in den Artifakt.
Ich verletze die Nummer, bitte.
Wir werden es später verletzt.
Und was wir sehen, ist das Nummer ist viel lower, meaning es gibt viel weniger Iterations an der CI-Level für die Unternehmen, die neigen Times-Einigerer sind.
Does it make sense what I'm saying?
So they are moving faster exactly because they have structured their work and their pipeline so that it flows.
There's not a lot of back.
I mean, like I'm thinking water pipes, right?
There's not turbulent flow, which is exactly the engineering term for it.
There's not like I add my change and it comes back and it comes back and it comes back.
They are able to get changes into production more effectively with fewer reiterations at the CI level.
And again, there's a lot of ways to do that.
And we know that they are doing all these things at Tempe said are really good.
Number one is the more that you can modularize your code back to things we already knew about software, the better...
Das ist alles gut, weil ich in meinem kleinen Bereich arbeiten und ihr in deinem kleinen Bereich arbeiten.
Und sie werden nicht mit einem anderen, weil wir das architecten oder designed in einer Appropriate Art.
Die andere Sache, wie wir vorhin haben, ist es, dass viele Menschen leften als möglich sind.
Es ist nicht wie die besten Unternehmen, die es nicht mehr...
Produce fewer bugs or have fewer mistakes.
I'm sure they produce just as many mistakes or maybe more than the average, but they're detecting and resolving those mistakes earlier in the process before it gets to the CI part of it.
Could it be if they manage to detect it earlier and also streamline the process so that not as many defects are?
Exactly.
So they also don't need so many manual interventions and they trust the process more.
So I guess if you don't trust it, you have to manually review everything and get slower.
And if it always hits some error correction loop, you wonder, hey, what's going on there?
Let's have a look at it.
Yeah.
Ich weiß nicht, dass diese Unternehmen für diese Unternehmen, aber das ist nicht Teil der Studie, aber auch in den nächsten paar Wochen, eine Industrie White Paper über AI Code Review ist, und ich und viele andere Leute sind co-authors.
Und wieder, was wir finden, independent von meinem Erfahrung mit CircleCI, ist, dass die Unternehmen, die Unternehmen, die Unternehmen, die Unternehmen, die Unternehmen, und die Unternehmen die schnellen sind, ist es genau weil sie die Harness- oder die Eval-Loop super-effectivität.
Und es ist nicht komplett aus dem Gelegenheit, aber es ist, instead, alles, was man kann, via ein Computer zu verändern.
Das ist ein bisschen was ich, was ich sage, zwei Dinge.
Either LLM, oder sogar besser, etwas deterministic.
Every time we discover a thing that slipped past, we engineer it back into the pipeline, as opposed to just checking it by a human over and over and over again.
And a really inspiring talk that I heard a couple of weeks ago, unfortunately it was private, was about software is not new in this idea.
Und so zwei Industrie sind wirklich sehr ähnlich.
Semiconductor Manufaktur und Pharmaceutical Manufaktur.
Ich will dir warum ich sage das.
Wir produzieren Chips und wir haben Produzieren Semiconductor Chips für Jahre.
Da ist ein Trillion Dollar Industrie, ich bin nicht lying, Trillion Dollar Industrie, called Semiconductor Test Equipment.
Wow.
Das ist nicht Produzieren, die Semiconductors.
Das ist Produzieren, die Equipment.
Das ist die Semiconductor.
Wie ist das Thema?
Und Produce It Reliably means, so these are biological things, like they can get infections and they can be broken in chemical or biological ways.
And like, it's human level dangerous.
It's mission critical to make sure that when we're trying, when we produce a drug or some substance at industry speed, that we don't kill people.
And so there's another.
Trillion Dollar Industry, that's about, I forget the phrasing of it, but essentially pharmaceutical test equipment.
So we produce this batch of, I don't know, what's the new fancy thing, the diet drug, Wegovi or whatever, like we produce a batch of this like, you know, substance and we need to validate at.
machine rate at industrial rate, whether that is the correct substance.
And again, there's a whole trillion dollar industry that is an engineering industry around validating that the pharmaceutical outputs are pure.
That's quite interesting.
Because when you talk about this, it sounds that these industries have developed one set of tooling, one which I can just buy to test something.
And when I think about software development, I just asked Claude, what do I need to set up my testing?
And it comes with, I think it was 50 tools, something like this, to have everything covered.
And wouldn't it be a great idea to have something just, here you go, that's it, that's the whole harness you need, the whole...
und es wird funktioniert.
Ich weiß, ob wir das in der Industrie werden wollen.
Ja, ich glaube, wir sind in der Industrie.
Ja, ich glaube, wir sind in der anderen Richtung.
Ich glaube, das ist auch.
Ich glaube, du war in der Richtung, dass diese disciplines sind eigentlich Engineering und wir sind jetzt lernen, wie wir in Software machen.
Ja, ja, ja, ja.
Ja, absolut.
Ja, ich bin...
Es gibt keine Software Engineering.
Es ist mehr und mehr Engineering every Jahr.
Es ist nicht inherently nicht Engineering, aber es ist nicht sogar falsch.
Es hat sich ein Kraft für viele Jahre.
Ich werde deine Frage in einem Moment answer.
Aber wir müssen auf das für einen Moment auf das Thema.
Für die meisten von humanen Geschichte, making textiles, die du mitnehmen, ist etwas, was manuell in deinem eigenen Familie gemacht.
Similarly, furniture.
Für die meisten der Humanität, was manuelles und manuelles, was manuelles, was manuelles und manuelles.
Vielen Dank.
Ich habe es nicht mehr als erstes gemacht, sondern in der Lage, dass es sich an das Stadte auf der Ebene geändert hat.
Wir sind in das gleiche das Stadte in Software.
Software ist es, hat es für ein paar Jahre, und mehr mit AI ist es, es ist ein Kraft.
Again, nicht falsch, das ist eine Art von Kraft.
Das ist eine Art von Kraft.
Das ist eine Art von Kraft.
Das ist eine Art von Kraft.
Das ist eine Art von Kraft.
Das ist eine Art von Kraft.
Das ist ziemlich interessant, weil ich jetzt über die sogenannte Software-Infektion, die sogenannte Software-Infektion ist, weil jetzt alle Leute benutzen LLMs, um die Web-Pagungen zu entwickeln.
So es ist zurück zu der Familie, wo es kraftiert ist, bevor industrialisierung.
Ich weiß nicht, dass es das ist.
Ja, ich wirklich like das.
Ja, ich wirklich like das Model.
Ja, okay, so now I'm a, this is going to be a great conference talk, which I'm not going to give it, but from craftsmanship to industrial, to mass-produced industrialization to industrial scale personalization.
Yes, so that's what it looks to me like now, but there's still the risk of wipe coatings that, yeah, you have the wrong harness.
So if we...
Ja, das war die Frage, die interessanteste Frage, die ich nicht answerte.
Ich denke, das ist es, und ich denke, wir werden, in der gleichen Weise, wie die Kompilers und Kompiler-Tool-Chains, die sehr bespoke sind, Very unique, very almost individual.
Different teams.
You and I are old enough to remember if you ever did any Windows development or Microsoft development, the first thing you do is you write your memory manager.
And then you write...
I hope there are some nodding people listening as well.
So you write your memory manager, then you write all your data structure implementations.
Who has a great team?
Nobody in their...
Nein, aber ich meine, nicht nur fünf Leute in der Industrie oder so sehr, sehr viele Leute, die auf die Möglichkeit sind, die auf die Möglichkeit sind, und ich respecte sie.
Und es gibt sehr viele Leute, aber ich bin nicht nur, dass sie existieren, dass sie auf die Struktur implementations sind.
Und wir haben über das in der Industrie in der Industrie, bei standardizing, ich meine, wenn du in .NET oder .NET, da ist ein Standard Set, ich weiß, du wissen, das ist ein Standard Set.
Ich weiß, du wissen, das ist ein Standard Set.
Das ist das, was ich schon gesagt habe, ist das, was ich schon gesagt habe, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich das jetzt, wie ich a world-class set of evaluations.
But I would have to assemble, like, we would have to assemble that.
And what you're saying is, I'd love to get that assembly out of the box.
And I agree.
Yes.
I still remember the times where everybody had to write his or her own string class because...
Es war besser, es war mehr performant.
Ja, so nobody does this anymore.
Ja, und jetzt, so, again, just to say, just to put the things together, I think, not I think, we are right now in the write-our-own-string-class phase of developing these evals or these harnesses.
Yes, everybody writes their own harness and says, yes, my harness works.
Ja, ja, okay, just to just to say that So, strong agree, and let me say that in more general terms.
Let's see.
There's a wonderful framework from Kent Back, which he calls the 3X framework.
It's explore, expand, extract.
Explore, meaning try a bunch of new things, expand, find one or two of those things, and really scale them.
And then extract is...
like taking profits in some sense, like using that as the baseline for the next thing.
So explore, expand, extract.
And we're in that explore phase, meaning we're, as an industry, we're trying a lot of different things to evaluate LLM output.
And the next phase, which we'll have all together, is the...
Standardizing around some small number of them and then leveraging them.
Similarly, in machine learning, there's explore-exploit.
So try a bunch of machine learning.
I've been doing machine learning for 25 years, whatever.
Doesn't make me an LLM expert, that's for sure.
But I do know a little bit about machine learning, as I know many of your listeners will.
And explore-extract.
We now talked about a lot of things where we do agree.
What we talked about with the harness and the abels, it sounds for me like a decision whether we do it by ourselves or wait for a product to hit the market and just flow with the market.
What do you think?
für ein large Unternehmen?
Just use the tools which are on the market and wait for them to get better.
I mean, each month they are getting better.
Yeah, that's a good question.
I don't think I have a good general answer, honestly, because this is very context dependent.
I will say, what's the way to say this?
Everybody who produces software with LLMs needs evals.
That's for sure.
Even without LLMs.
But particularly, again, back to our story about LLMs are forcing us to do good software.
We already knew how to do good software and now we are absolutely forced to do good software if we're going to leverage things at machine speed with LLMs.
I think it is very, it is, I think right now, Ich könnte und ich könnte, und ich könnte, und ich könnte, eine Firma, könnte ein sette E-VALs mit ein paar off-the-shelf Dinge, wenn ich sie von einem Vendor oder eine Open-Source kaufe.
Ich denke, wir können alle das tun.
Ich denke, ich denke, wenn wir zu tun, ich bin überzeugt, wenn ich mich, wenn ich LLM zu produzieren, in meinem Unternehmen habe ich auch, wenn ich ein LLM zu verwenden, ich muss auch, wenn ich ein LLM zu verwenden, das ist ein LLM zu verwenden.
Das ist einfach ein Fakt.
And I think you can't have both.
You can't say, I want to produce code at LM speed and I'm going to wait to evaluate it.
Right?
Yeah, yeah, yeah.
But that's why I get a feeling many companies still have the foot on the brake because they say every line of code has to be reviewed manually and because they don't have the evals and even if they have, they don't trust them.
Well, I'm sure that, well, that is true.
That's absolutely true.
And what I, I'm going to answer that in a second.
The overall, I don't want to give people the wrong impression.
The overall industry has a very wide distribution of AI usage, very wide.
So there are those companies that I mentioned that are 9x, do have 9x, 9x faster, fewer merge conflicts, more flow than even the average.
They are on the, you know, big users.
Ja.
Die Individual developers, wir haben diese massive und individuelle Team.
Wir haben viele Teams, die wir als Software Factory haben.
Und sie sind sehr klein und sie benutzen AI und evals.
Und sie sind generatet, mostly Greenfield, vieles von neuen, gutem Arbeit.
Und dann, noch in unserem Unternehmen, CircleCI, die meisten von engineers sind in der Arbeit, für sicher.
but are still working on building the eval harnesses that we're talking about.
So again, I want everybody to just like not feel bad if you're in the large part of the distribution rather than this small part.
But this part's growing.
Okay, so you're asking what should a company do?
And is it, what's the way to say that?
You're asking, I'll frame your question differently.
How can we use Leverage LLMs to build at machine speed and evaluate that building at human speed.
You can't.
I'm not saying give up, but you can't do that.
So what should you?
Well, what we used to do is we built at human speed and we evaluated at human speed.
That worked, kind of.
Aber du kannst nicht auf machine speed und evaluieren at human speed.
So, was du musst, all die Dinge wir sprechen über, ist, dass du die Humans, dass, stattdessen, dass jeder Human evaluiert, jeder Einwohner, jeder Einwohner, und in der Hilfe der Dinge, die in der Industrie, eine Pipeline, eine Eval-Set, oder eine Harness, um das Wort, um das zu evaluieren.
So, what I'm not saying is, and then walk away and it's machines all the way down.
What I am saying is engineer the harness exactly like semiconductor test equipment manufacturer, exactly like pharmaceutical equipment manufacturer.
Do that.
And then you will have machine production of code and some amount of machine evaluation of code.
Now, what's left, that is okay for humans to evaluate.
Does that make sense?
Yeah, it makes sense.
So in my own words, I would state something like, it makes sense to have a team of toolsmiths who build a harness in order to enable the rest of the developers to go faster because they can rely on this harness, on the evaluations.
I think not every developer should.
und dann kann man die Türen sehen, wenn sie die Türen sehen können.
Man muss die Maschinen speed für die Reviews und ja, wie man muss, und dann muss man es und entscheiden, wie.
Ja, sorry to interrupt you.
I get excited.
No problem.
We're finally agreeing.
And one of the users typed in, the engineering challenge moves from code to harness.
Beautiful.
Wow.
Couldn't have said it better myself.
Absolutely beautiful.
And that's what we're trying to say, is you can't, again, it's just a sayback, you can't do machine speed code production and human speed code evaluation.
Those don't work.
Und das ist nicht, dass wir nicht mehr, dass wir uns ein bisschen mehr in den Abstraction haben, und die Evaluations haben, und die Specs-Up-Front haben, aber die ganze Zeit, und zu Ihrer Meinung, die ich sehr gut gemacht habe, ist das jeder, dass das jeder Engländer von 1,000 oder 10,000 ist?
Nein, das ist nicht für andere Reiseinschaft.
Nicht weil sie nicht tun können, aber Das ist ein Duplicative von Effort und das hat sich selbst Probleme gemacht.
So, every company, an large scale, diese Zeit, das ich bin aware, das ist das gut, hat eine Art von Platform Team, Developer Experience.
Du kannst viele verschiedene Worte für das.
Ich ran das Team an Ebay, als ich die Chefarchiteite war.
That is the team that engineers the production of code for the company.
By the way, I know everybody knows this, but I'm going to say it out loud.
Before AI, we already needed that team.
And now we really, really, really, really need that team.
Or we need the help of that team, right?
Again, just like...
LLMs are forcing us to do good software development techniques that we've already known.
Similarly, LLMs are forcing us to do good, I don't know, software development organizational techniques.
Like, okay, if everybody's doing the same thing or roughly the same thing, let's extract that out into a common service or a common team or a common platform and do that.
That's just...
Das ist einfach nur Efficiency.
Aber die Problem, ich sehe, ist das Engineering ist fun.
Wenn Engineering ist jetzt mehr in der Harness als in der Coding, dann wird das Harness nicht produziert.
In der Terme von hier ist ein Produkt, das ich kann.
Es produziert Wert in der Terme von hier ist ein Produkt, das ich kann.
Es produziert Wert in Terme von...
Yes, hier ist mehr quality.
Like the test system in the semiconductor industry, where you only have a few companies who produce it and everybody uses it.
And the value is in the different types of semiconductors.
And so I wonder, I think we will not need so many harness engineers.
So the fun jobs.
In Hahn's Engineering, where not be so many?
Oh, I don't know.
I have several thoughts.
There's a thing called the Jevons Paradox.
Yes.
And you're familiar, but just to say briefly for those who are listening that aren't, Jevons was an English economist in the mid-1800s, and he was talking about coal.
Aber was er hat er, was sehr surprising, und das ist der Paradox, ist, dass wenn die Kohlrekturkampen wurde, es war cheaper zu bekommen, es war cheaper zu bekommen.
Somehow, in aggregation, Britain spent mehr auf Kohl.
Wait, was?
Du hast es cheaper zu bekommen, und yet in aggregation, wir spent mehr?
Warum?
Weil Kohl war 10x cheaper.
Jetzt sind es, was nicht so, was Kohl war.
Jetzt sind es.
Das ist mit dem Kohl.
So, jetzt sind wir.
So, ich weiß, was ich.
So, ich weiß, was ich.
So, ich weiß, was ich.
Ich weiß, was ich.
So, ich weiß, was ich.
Ich weiß, was ich.
Mit LLMs, wir haben die Kosten von einem Teil des Software zu produzieren.
Das hat sich schon viel cheaper.
So eine Partie ist, die engineering Probleme verabschieden und auf dem Upstream.
Was sind wir, was wir bauen?
Was ist das?
Und wie wir eine Speck verabschieden, das?
Und dann geht es um die Engineering, die die LLM-Maschie verabschieden ist.
Das ist eine force.
Das ist das, was wir tun?
A resource gets substantially cheaper.
The other part is that the demand for the resource is what's called elastic.
In other words, there's no limit to the demand for this thing.
And so like energy, there's no, sadly for the planet, there's no limit to our demand for energy.
There is a limit to our demand for food, for example.
Like there's only so much food I can eat.
But there's, but.
und Software ist kein Problem für die Software.
Es gibt immer mehr Dinge zu automatisieren, mehr Dinge zu machen, mehr Dinge zu machen, mehr Dinge zu machen.
So, als LLMs, die Zeit der Zeit, die Zeit der Software-Infektion verändern, haben wir viele sehr wichtige Aufgaben, und auch, manchmal mehr Dinge, wir werden mit Software-Infektion machen.
So I'm not scared for our profession.
I think there will be more engineers tomorrow and the next year and the next year than there ever were.
Sounds good.
And yes, I think this comes also again down to am I a programmer?
Do I just type the lines of code or am I a developer who solves problems?
And we will have problems also in the future.
Das ist wunderbar.
Ich habe das in einem anderen Fordergrund geplant.
Ich liebe das Programmer-Developer-Win.
Das ist wirklich gut.
Ein anderer ist, sind Sie die Reise oder die Reise?
Oder sind Sie die Reise?
Und das ist nicht besser als die anderen.
Es ist okay.
Es ist okay zu haben, die Typen.
Ich wirklich lief die Typen.
Es ist wirklich fun.
Es ist Problem-Solving.
Und dann die andere Related-One...
Ich denke, ich habe etwas falsch mit dem Mikrofon.
Oh, ich bin sorry.
Kann ich mich jetzt?
Ja, jetzt sind es ein disturbances.
Okay, ich bin jetzt nicht mehr.
Wir sind jetzt nicht mehr.
Aber, um, der andere Distinction, die du raisest, ist Craftsperson versus being an Industrie engineer.
Ja.
So, we now agreed on lots of things and we are already shortly over time.
I would like to go back to the start where you said, yes, it's AI, but it's not like human intelligence.
And we are in an industry where we invented duck typing.
If it walks like a duck, walks like a duck, swims like a duck, it probably is a duck.
Wenn es jemand intelligenter antwortet, hat einen Grunde und intelligenter antwortet, was du missst mit AI-Systemen?
Wo du die line ausstrahlst?
Oh, wow.
Ich weiß nicht, dass ich etwas Besonderes habe.
Ich meine, wir haben diese Frage.
Since Marvin Minsky in the 1950s, what does it mean to be a human?
And the Turing test, of course, even earlier than that.
I don't want to say it doesn't matter, but for the purposes of this conversation, it kind of doesn't matter if that makes any sense, right?
So like using LLMs as tools in software engineering, they are very helpful.
There's just no argument about that.
They can do a lot of things.
Und wir sollten es ein neues Tool machen, einfach wie ein Compiler.
Wir haben jetzt Dinge, die wir uns haben, haben wir, und wir können die Maschine helfen.
Die Compilers haben nicht gemacht, die wir nicht machen, die wir nicht machen.
Die Compilers haben nicht gemacht, die wir nicht machen, die wir nicht machen.
Und Kobol wird, die wir nicht machen, weil wir nicht haben, die wir nicht machen.
Wir haben uns nicht gemacht, die wir nicht machen.
Und wir haben das nicht gemacht.
Ich weiß, es ist ein super interessantes Thema, als ein Mensch, und von einem Philosophischen Perspektiven.
In der bestimmten Software Engineeringen, ich finde, dass es das was die Tool macht.
Was soll ich sagen?
Ich finde, ich bin, ich bin, ich bin, ich bin, dass es ein Mensch, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, ich bin, Und wieder, ich habe keine besonderen Insight into das, zu sagen.
Aber ich denke, die Geschichte ist, ich habe immer noch einen Anlauf.
Die Geschichte ist, die wir uns nicht verstehen, sehr gut.
Und jederzeit wir denken, oh, das wird es zu werden, wird es nicht.
Wir haben nur noch mehr.
Da sind wir noch mehr.
So ich denke, es ist eine andere Sache.
Ja.
Quite interesting.
I also wouldn't compare it to humans, because humans are even more than just intelligence.
But I think it's quite close to, I mean, it's a different form of intelligence, but quite close to the definition of intelligence we had during the last years.
Yeah, and I think what we're learning is our definition of intelligence.
Our definition of what it means to be human is evolving.
Every time we figure out how to automate or industrialize a part of what it means to be human, we find new things that are unique about humans.
I don't know.
Last question before we end this.
Do you think AI and LLM can be creative?
Sure.
Yes.
Yes.
Well, no, no, I don't know.
Maybe I'm slow because I'm a human.
Yes, if only because combining existing things in new ways is creation.
How do you know that?
There was a time in my life where I did patent law briefly.
We will not go into that time in my life except for later, some other time.
But it is totally...
Patents are about what's invention.
And you can have a lot of arguments about whether the definitions there are good, and I have them.
But combining an idea from one discipline with an idea of another discipline in a new way, that is creation.
That's how humans do it.
Ich denke, LLMs können sie werden.
Sie können sie alle ein bisschen von creation machen.
Das ist ein Philosophie.
Ich bin nicht mehr oder besser als jemand.
Ich habe noch nie gesehen, oder songs oder Musik.
Die LLM-produkte ones sind einfach.
Es ist ein bisschen besser.
Produzieren etwas, das ist wirklich gut.
Aber Produzieren etwas, das ist wirklich gut.
Aber Produzieren etwas, das ist wirklich gut.
Aber Produzieren etwas, das ist eine neue Art.
Artistic creations.
Maybe I'm making a distinction between creative and artistic.
How about that?
Absolutely creative.
Yet to see whether it's artistic.
I don't know.
Sounds good.
So we are already quite over time.
Thank you for your time, for your insights.
And I'm really looking forward to meet you for your...
You have a keynote at the Software Architecture Gathering.
So it's...
The title is something like, You Can't Rewrite It All, about modelization.
Yeah, so it's a lot of the things we talked about.
So very briefly, the keynote was inspired by the so-called SaaSpocalypse, so the software as a service companies for several months at the beginning of this year, I believe.
Yeah, this year.
Something like $300 billion of enterprise value was removed from Salesforce, SAP, blah, blah, blah, because everybody thought nobody needs them anymore.
We're all going to rewrite those pieces of software ourselves.
Anybody who really worked on them knows how insane that was.
And we're seeing that the market has recovered.
But I think that taught us a lot.
And we talked about a lot of these ideas here where the power of LLMs is teaching us.
Modularity still matters.
Evaluations still matter.
Invariance and architecture still matter.
In fact, they matter even more.
And so you can't rewrite it all is my idea of it might seem like I could just regenerate a customer relationship management system or a database or an airline reservation system by myself with my trustee LLM.
I don't think that's true.
Und die Talk explores why das ist true und also die Ways-Forward, wenn das macht.
Ja, ich bin wirklich glücklich.
Ich habe bereits gebrochen die Konferenz slot, so ich werde es beinwerte.
Danke für Ihre Zeit und danke für alle, alle, die in diesem Video zu sehen.
Ja, haben wir eine schöne Woche.
Danke.
Bye.
Hi, I am Alarhoch und Bach.
Do you organize any user groups, conferences or other tech events?
Then feel free to add them to Trev.bundesch, an uncommercial platform for tech events in the German-speaking community.
It's free without any advertising or tracking.
Just visit Trev.bundesch.
www.trev.tesch.de
