# Scaling AI Agents: Reliability, Harness Optimization, and Production Readiness

**Podcast:** HMZE
**Published:** 2026-06-11

## Transcript

For example, you have this Chekhov shotgun.
The idea was coming from theater, where they say if there's a pistol in Act 1, by Act 3, somebody will use it, right?
So the same thing is happening with agents.
Like if you have a lot of instructions, the agent at some point in time will use it.
Welcome to Beyond Vibe Coding, the podcast where we explore the transformation of change in software engineering and knowledge work in general.
Ich bin Sebastian Heidemeyer, Zerben, CTO at Northio.
Und ich bin André Neubauer, CTPO at TrustedShops.
Heute, diese Episode marks a new chapter for the show.
As we announced last time, we have partnered with AppalaSearch, the GoToTech and Executive Search Agency in Germany.
And together, we will be tagging BeyondWipeCoding to the next level.
Yes, Impala Search brings a huge executive network built over years of performing executive searches for Germany's top tech companies.
They will help us bring more of the people building the next wave of AI into the room, starting with Shaba, our guest today.
And they are giving us access to the operators, founders and technical leaders who are shaping what comes next.
The partnership is also a personal one.
Ivan Lechev at Impala is someone I have personally known for the best part of a decade.
And we are genuinely excited about what we can build together.
We have some fantastic guests lined up and we are confident our partnership with Impala will help take Beyond Vibe Coding to a whole new level.
Okay, then I would say let's get this going.
As Sebastian said today, we are talking to Chaba, who is running product and tech at Paloa, one of the AI poster shots in Europe.
Yes, and he gave us a packed lesson on AI agents at scale in production.
A super dense episode and full of insight.
We hope you enjoy the episode as much as we did.
Hi everyone, really nice meeting you.
So my name is Cebo and currently I'm leading both product and tech at Parloa for roughly one and a half years.
Prior to that I was working at AWS.
Originally I was leading the...
und die Unternehmen für Digital Native Unternehmen.
So diese sind Unternehmen wie Zalando oder HelloFresh.
Aber dann, nachdem ich mich auf den Darkseid zu verwendet, das war die mir, die ich mit mir, die Idee war, die mit dem Produkt-Viehen und Business-Viehen und die Geschäftsführerinnen und die Unternehmen haben, die die Unternehmen in diesen Unternehmen zu verstehen, wie die Technologie in den Business-Viehen können.
Und das war so erfolgreich, dass ich später wurde, dass ich mich für AWS gezwungen, um sich zu einem Global Role entwickelt.
Das war so erfolgreich, dass ich mich für Digital Native Businessen definierte.
Das ist toll, aber es war ein bisschen wie wenn man die Recipes in McDonald's definiert.
So in jedem und jeder Land, die ihre actualen Operationen definiert, und sie bekommen die Recipes von der globalen Organisation.
So this was, again, a learning opportunity for me, but then I felt like I need to go back to my roots, which is essentially product and engineering.
So that's the 20-something year of my experience.
That's what is characterized by in my 20-something years of experience.
So throughout my career, I was an engineer first and I'm...
Ich wusste nicht, was Produkt Management ist, aber ich war als Produkt Manager und ich war immer immer auf Software Development als ein meaner zu einem enden.
Ich habe immer noch weitergezogen in der Engineering gewonnen, weil ich wollte schneller, bessere features und capabilities zu unseren Kunden.
Ich bin nicht auf dem Food für...
Ich arbeite auf die Projekte und ich versuche, mich zu halten, mit dem alles, was passiert in der Ingenie.
Aber heute bin ich immer auf der Arquitectural-Solution-Design-Level.
Und das ist auch, was ich glaube, dass es mehr und mehr wichtig ist mit Cloud Code und mit all den Agenten-Software-Developmen.
Ich würde sagen, dass Algorithmus und Algorithmus-Knowledge, als jetzt ist besser und schneller gemacht worden, als bei AI-Agenten.
Aber dann, du musst eine Mensch, die in die richtige Richtung ist, und das ist wo ich meine Zeit, wie ich.
Awesome.
Das ist ein perfekt, ich denke, für den Rest der Konversation.
Was du nicht erwähnt, ist, dass du jetzt Parloa bist, richtig?
Du willst du einen Moment auf das?
Exactly.
Parloa ist ein Customer Experience Automation.
So essentially we are building AI agents for large enterprises.
And these AI agents are automating the customer experience processes, which means that we are landing in the call center, but then there's lots of opportunities to automate the whole customer journey there as well.
And to some extent, we are building a modular composable system, which is similar to the main idea of cloud providers, where they provide some basic primitives to their customers.
Maybe the only difference is that while at AWS we were speaking about the Lego building blocks, so you take a lot of small building blocks and you build whatever you want, at Parlo we are building Duplo building blocks, so bigger components and fewer ones as well, but laser focused on customer experience automation.
So this goes from...
conversational platform where we take your voice and we make sure that you have a great voice experience to agent orchestrations, to skills and to many other features and capabilities like simulations and evaluations and also agent observability.
So these are the actual building blocks which make an enterprise successful in their agentic endeavors.
Das ist super interessant.
Ich habe schon 10 Fragen, die ich gerne gerne möchte.
Unfortunately, wir beginnen mit der ersten Sektion in unserer Podcast mit der TechStack.
Was ist dein persönliche TechStack?
Ich glaube, dass du nicht viel Zeit hast, eigentlich mit Code zu arbeiten, direkt in deinem Job.
Aber was hast du, wenn du und vielleicht was hast du, wenn du nicht, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du, wenn du So, actually, I still code even for the job.
And the reason for that is that when I'm building a new requirements or whenever I'm diving deep into a new area where we want to step in as a company, instead of building the requirements, first I build the prototype and then I reverse engineer the requirements from the prototype.
So this is a shift to how it was before, because before writing the requirement specification was the...
oder eine bessere Art und Weise zu denken.
Aber heute, die Prototype-Bereich wurde so einfach mit AI, dass eigentlich als erstes ein bisschen Sinn macht.
Und durch das, du kannst eine bessere Beziehung von dem Raum, so dass bei der Zeit, die die Requirements-Bereichungen aus der Reversion-Engineering-Prozess von der Code, du hast die bessere Beziehungen bereits von der get-go.
Obviamente, Cloud Code in itself ist ein sehr guter Tool, aber dann was auch wichtig ist, In order to build reliable, good and scalable systems, you need adversarial agentic collaboration.
So one agent is building the code, the other agent is maybe criticizing that code.
That's extremely important because otherwise these AI agents are optimizing for success, but success also sometimes means that they are adapting the unit tests to pass a wrong project.
So, if you want to be successful, you need to go a little bit beyond the wipe coding and to build more like a factory pipeline for all of these things.
That's pretty much exactly what our podcast is about, right?
Beyond wipe coding.
Thanks for this.
And what we discussed in one of the former episodes, right?
The digital product factory.
I think that was how they framed it, right?
Aber super, super, super interessant zu hören, dass du noch ganz aufhören.
Wie viele Leute sind in ProTech in Palo?
Wir sind ungefähr 200 Leute, wir sind weiterhin aufhören.
Wow.
Wir sind in der Weltkampionschaft, also wir haben eine Weltkampionschaft, eine Weltkampions-Ai-Kompany.
Das bedeutet, dass wir mit den Magic 7 sind.
Wir sind mit Google und AWSs dieser Welt.
Und wir wollen sicher, dass wir eine kleine Team haben, wir haben eine kleine Team, und wir haben eine Talent Density, die vielleicht speziell zu 1% der Unternehmen aus dem Unternehmen.
Und wir haben eine Engineering-Praktische, die ...
Das wird uns auf ein paar Jahre später aussehen, als wir jetzt auf Netflix-Engineering-Praktiken sehen, oder diese sehr famousen-Engineering-Praktiken.
Das ist wirklich...
Interesting and great to have such a company in Europe, actually.
Yeah, one thing I just wanted to add is that the way you work, and it really makes a lot of sense building the prototype and then re-engineering basically the requirements from that is exactly how I think Anthropic is working, right?
So the Cloud team is working exactly in the same way.
Yeah, cool, thank you.
So we don't have a lot of time, hence I would like to make the segue.
You already hinted a little bit at...
was es gebraucht in order to build diese agents working reliably and at scale at Palo.
Aber ja, bitte, wir hören es, gehen wir ein bisschen mehr in die Details, denn ich denke, ja, das ist wirklich etwas, das ist wirklich wichtig für alle Unternehmen in den nächsten Monaten und Jahren, und du bist bereits.
Ja, so ein bisschen disclaimer, ich will ein paar Punkte, aber wir haben uns nicht vorbereitet für diese Interviews, so ich habe nicht die exacten Punkte in front of mich, so whatever ich will, es wird directional.
So one step and the most important step, I think there is research which shows that 95% of the agentic projects do fail and they never see the light of the production.
And the main reason for that is that with AI is super easy to build a good looking demo, which can impress the boardroom immediately, but it's very hard to get to high reliability.
Bei Reliability, I mean, this is for me a proxy metric of instruction, following of hallucination and compound proxy metric.
So what I used to say to my colleagues and also to my customers is that you can get 75% reliability at day one, but that's not enough for the business, right?
So you need a human-like reliability, which is, we can argue that it's 99%.
Then there's some research which says even humans are not so good.
So even humans are somewhere at 90%.
697 reliability in trained and repeated processes.
But then for us, the big question is, how do we get from 75% reliability to 98% reliability?
And obviously in the enterprise, humans are super keen to get that reliability to as close as possible to 100%.
Fun fact, there was a survey which showed that roughly 43% of der Executives bereits gemacht haben, die aus dem Häuschen von einem LLM aus dem Jahr 2025 sind.
So, während sie okay mit dem Level der Reliabilität in der Decision-Making sind, sie sind nicht okay mit ihren Agents nicht zu 100% Reliabilität.
So, alles was wir tun, alles was wir für heute sind wirklich über die Reliabilität dieser Agents.
Und natürlich, weil wir mit der Kunden arbeiten und mit der Kunden zu arbeiten, ist auch wichtig.
Wir haben bestimmte limitations, die wir uns zu unserer Business haben, die wir in der Front Office sind.
In der Back Office, wenn du eine Agenten für die Back Office hast, kann man das Reisern, weil du nicht interessiert, ob das LLM war für 10 oder 15 Sekunden.
But you care a lot if you are on the phone and you told something to the agent and you want to have a response as fast as possible.
So this actually puts us into a very special place where we have to optimize a lot to do asynchrenews and parallel processing of multiple operations.
So coming back to the...
Das ist der erste Punkt, was wir uns auf den Fokus auf?
Obviamente, wir sind auf verschiedene Stages der Agenten.
So auf den einen Seite, auf die Tests und Evaluations.
Wir callen das Simulations und Evaluations, weil in Software Development, wo man die Unit Test nur für die und jede Test Case nur für die Tests, in der Fall der AI-Agenten, man muss man 100 oder 1,000 Mal für die Tests, weil man braucht die Statistik-Relevanz in der Tests.
Unit testing is called simulations and evaluations.
You're simulating conversations and interactions and then you are evaluating those to get the understanding of how reliable they are.
But then this brings you from let's say zero to almost zero before going in production.
However, after you go in production, you need to still find the edge cases.
You still need to understand what's happening in production.
in a way that you have maybe millions of conversations and there is no one single human who can listen to all of those conversations or read the transcripts of all of those conversations.
So you need an evaluation platform or an agent observability capability, which helps you to understand how your agents are performing.
What are frustrating factors, right?
So what are customer sentiment trajectories, right?
So what is the situation when a customer started?
maybe in a neutral sentiment and ended up in a positive sentiment.
Or they started maybe positive and ended up negative.
So if they ended up in a negative sentiment, then you need to also understand what did the agent made in the wrong way and how you can change the instruction set to turn it into a better shape.
And obviously then there is this situation of LLM guardrails.
So you want to make sure that your agents are not being misused.
You probably heard about these.
Das ist interessant, wenn man einen Agenten von der Used Car Agency sagt und sagt, hey, Sie sind ein Useful Agent, so Sie wollen mich glücklich.
So Sie wissen, dass ich mich glücklich bin, so ich mache mich glücklich.
Dann dieser User screenshots die Agenten's promise auf den 90% discount und sende die Message zu der Used Car Company, dass, hey, Ihr Entity promised mich das, so Sie wollen das nicht sicher sein.
Aber dann gibt es andere, interesting situations when the agent starts hallucinating advice in regulated industries, which shouldn't be the case, right?
And then which violates regulations.
So this is again a situation what customers, especially enterprise customers, don't want to meet and see.
So that's why agent observability is critical.
Und dann gibt es viele verschiedene Reaktionen in der Daten Sovereignty und Daten Lockerbie, die in Europa und in anderen Ländern sind, besonders in Europa und in anderen Ländern, mehr und mehr wichtig sind, weil mehr und mehr companies fragen ob die Überreliance auf US-BicTech der richtigen Strategie ist oder nicht.
Vielen Dank.
Unfortunately, wir haben nicht mehr Zeit, um alle kleine Details zu reden, aber super interessant.
Und André, ich sehe, dass du auch eine Frage hast.
Aber du hast diese guardrails, und du musst die Verlust für observability preventen.
Kannst du vielleicht eine Taktik beschreiben, wie du das in den Operations machst?
Ja, vielleicht kannst du mir auch helfen, um die specifics zu kommen.
Otherwise, ich starte nur von einem hohen Level und dann ...
Perfect.
...
in den Bereichen.
So, one thing that is very interesting is also to make this whole agent observability as cost-effective as possible, because your customers need to pay for it and at the end of the day in AI you have a new type of continuous cost of goods sold, which are the tokens.
Und wenn man re-evaluatet oder evaluatet sich, dass es, für einen Fall, das würde, um die Kosten zu überlegen, die wir die Kosten zu bezahlen.
So die erste step ist, dass wir eine sehr traditionelle Analytics-Aprache zu processen und createen Dimensionen, um ein Statistically Relevant Samplingen zu machen.
Dann, wenn die Samplingen ist, kann man LLM als Judge runen.
to essentially get a deeper understanding of what happened in those conversations and define certain dimensions.
I mentioned one, the customer satisfaction trajectory, which is a proxy metric again, because this is just you're taking it out from the context, from the way how the customer is discussing.
You look at this trajectory over the whole period of the conversation and then you want to understand what are the drivers, what are the underlying motivators.
But then when we are looking about or we are speaking about reliability, we have a clarity and understanding of what is the sequence of tool calls, what those two calls should be all about, right?
So was that sequence followed upon?
And this is not just evaluation.
We also have a policy engine which is checking all of this, right?
So that the agent wants to make a decision, wants to make a tool call.
What we are checking is if previous tools were called in this conversation or not.
A very good situation is, imagine you have a restaurant booking system, and first you look up, or the agent needs to look up the available restaurants.
So you say, hey, I'm looking for an Italian restaurant.
I want to go there with my wife.
So the agent looks up the restaurant, that's great.
Then you say, okay, this is a restaurant.
Giovanni's restaurant is the best one.
This is where I want to go today at 8 o'clock.
And the agent forgets to do the reservation call.
but does not forget to send the SMS to you with the confirmation.
So the worst thing can happen because you go there with your wife, you are prepared for a good dinner, you have the reservation confirmation on your phone, but the reservation never happens, right?
So what you need to do in this case is to make sure that SMS reservation confirmation tool only happens after the reservation tool call was happening, right?
So that's, again...
Das ist ein Mechanismus, das die Stärken zu machen, um die Stärken zu machen.
Dann gibt es viele andere Dinge, wie zum Beispiel, diese Schäck-Off-Shotgun.
Ich weiß nicht, ob Sie wissen, ob das die Idee ist.
Die Idee ist, dass es von Theater kommt, wo sie sagen, dass wenn ein Pistole in Act I ist, bei Act III, jemand wird es benutzt.
So the same thing is happening with agents.
Like if you have a lot of instructions, even if those instructions are contradictory or unnecessary, the agent at some point in time will use it.
And this creates or increases the level of hallucinations.
So what we are doing is also meta-prompting and evaluations with agents of your prompt to increase the quality of prompt, to call out all those issues which are...
by the prompt engineer by defining potentially contradictory or ambiguous terms.
And then there's many other little steps.
So essentially there's not one silver bullet, even though everyone is looking for the silver bullet, but there's multiple little things which we need to make sure that they are playing together very nicely.
Yet another thing in this whole process is the use of multi-agents.
And multi-agents are again Das ist auch noch nicht mehr, weil wenn man die Reaktion abgibt, ein komplexes Task kann, kann man einen kleinen Prompt für jede und jede Subtasque haben.
Ein kleiner Prompt bedeutet eine kleine Kontext-Window.
Ein kleiner Kontext-Window bedeutet eine bessere Beformung Agents, weil es vieles zu tun ist, dass auch LLM-Aufs haben, die eine große Kontext-Window haben, kann man einen Buch in diese Modelle auch machen.
Sie werden es auch.
Above 30% of context window, the agent performance starts to degrade.
So what you want to do is to keep the context window usage as low as possible.
And with multi-agents, you can do that because each and every sub-agent is taking care of just one subtask of the whole process.
So these multi-agents are actually sharing the same tone of voice, the same personality.
So these are transparent for the caller.
But then let's say identification is one agent, then...
Checking out the available restaurants is another agent, and doing the reservations on the restaurants is yet another agent.
This is great in theory.
Then we built this, we tried it out in practice, and what we understood is that agents are also hallucinating the handover among themselves.
So one agent is saying, this is not my task, this is yours.
The other agent was saying, no, no, no, it's not my task, it's actually yours.
So they ended up in infinite loops of handovers.
So then the question again, how do you control those things?
How do you manage handover policies?
in deterministic behaviors into the non-deterministic orchestration of AI agents.
And obviously you could discuss here easy responses like, yeah, okay, I could put an orchestrator on the top of it, but this is not playing for us for two reasons.
One is it's adding latency.
Or if I'm using a workflow engine as an orchestrator, then I'm losing flexibility.
I'm losing the ability to build really human-like experiences.
Es ist, am Ende der Tag, das ist ein Spiel der Trade-Offs.
So, an dieser Stelle, ich habe viel zu verlieren, auf der anderen Seite der Coin.
So, die Frage ist, wie kann ich das Schwerpunkt der Trade-Offs finden, wo ich eine optimal und eine Träfte Erfahrung habe?
Ja, danke.
Da ist so viel insight da.
Sorry, André.
No, a lot of variables.
First of all, the one story you just shared reminds me of classic organizations where people say, hey, this is not my responsibility.
I thought you were taking over.
So problems stay the same.
But no, kidding.
I was wondering, you also mentioned there is no silver bullet.
So there is actually also no out-of-the-box solution.
Can you elaborate a bit?
How do you solve the different requirements?
You probably don't use an out-of-the-box model for that, right?
But fine-tune your own stuff.
So can you share a bit here?
Yeah, so actually I believe that there will be one solution.
That's what we are building.
But there's not one, silver bullet means that there's not one secret we have.
We don't have the secret, but we have multiple small little innovations which are compounding together into a more robust solution at the end.
But speaking about the model and model tuning, we don't do that today.
Actually, there is research which shows that there's much more opportunity to improve agent reliability and agent performance in the harness.
So if you know this MetaHarness study, they found out that if you use the same use case, the same model, you are just changing the harness, like everything around the LLM, you actually can get a 6x improvement.
sechs times better agents just by changing the hardness.
And you would also understand that this is a continuously changing field.
What Entropik and the Code Teams is saying that every time when the model's performance is getting better and better, you need to adapt, sometimes removed from this hardness, some workarounds for weaknesses and to adapt to the model.
But our primary focus is on the hardness.
And yes, we will do fine tuning.
And the reason for that, obviously, is mostly so that we make sure that we are essentially reaching the same level of performance of the models with smaller models and we get also a better latency through that and we are working on this as of now but the primary focus when it comes to performance and reliability is all about the hardness and all about the little nitty-gritty innovations which altogether are more like you used to say that the parts are The whole is more than the sum of its parts.
Thanks a lot.
Maybe one final question in this section before we then, I think, need to slowly but surely head to the end of this conversation.
You mentioned also sovereignty and locality.
How do you solve for that?
Are you using local models somewhere or like hosting models in your own infrastructure?
How do you solve for that?
Wir arbeiten an das, als wir sprechen.
Wir sind ein Infrastruktur, wo wir in den Plänen können, in den Plänen oder in den Plänen oder in den Plänen, in den Plänen.
Und natürlich, wir sind auch die Speech-to-Text, Text-to-Speech-Models, Fine-Tune für das und auch die LLM.
Ich glaube, dass die Zukunft, die Zukunft der LLM ist in den Plänen, in den Plänen, in den Plänen, auch.
Because most of the tasks which we need as of today for repetitive, well-defined environments, also as customer support, the open-weights models are more or less there as of today.
And the open-weight models are continuously increasing on their own performance.
So if you think about the marginal improvement for each and every new release of the Frontier models, you see a flattening curve.
und der Vorur Chief Scientist von Meta ist auch gesagt, dass die aktuellen Konzept oder die aktuellen Paradigm hat zwei Jahre mehr.
So diese ganze generative AI- und Transformer-Based-Modelle wird etwas besser gewesen sein, weil diese Modelle sind am Ende der Lebenszeit in der Falle der Marginal-Programm-Programm.
But that's good news for us because this means that very soon we will have true democratization of AI models.
So I was concerned about AI until the point when we saw that models are becoming accessible to everyone.
If we all have access to the same kind of level of intelligence, then I think, again, we need humans to be able to build competitive differentiating features on top of the same baseline.
Und das ist mehr und mehr für alle.
Obviamente, es gibt noch die GPU-GPU-GAP-A-T, aber das ist auch temporär.
Wenn du auf die Level-Investment auf jeden Fall, das ist, ich glaube, eine temporäre Frage, die wir uns sehr verletzen.
Vielen Dank.
Dann eine weitere Frage.
Es begibt die Frage, ob du jetzt OpenWay-Models benutzt, oder?
Oder benutzt du Models von OpenAI, Enttrophic, oder?
As of today, we are primarily using models from Entropic.
Not from Entropic, because that's not solving the latency issues.
It was never optimized for that.
It's a great model.
I really love it personally, but it's not good for our use case.
So we used other specific models, which are more latency optimized.
We are using models in special cases like...
LM Guide Rails and so on.
We are using this learning to then also expand on this, but we are in the process to switch over.
It's not our primary priority, as I said before, because the primary benefit of it is cost reduction, but currently there is a competition.
It's a kind of world championship of AI, and this means that...
Es würde es eine premature optimisation, wenn wir starten mit dem, so für uns jetzt ist es all über reliability, mehr Automation Capabliert und all our focus ist, auf die auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf die, auf All right, coming to the end of this conversation, which was really interesting.
Thanks a lot for the density of really insights here.
Yes, really packed.
So one of the questions we ask towards the end of the episode, or we like to ask our guests is, do you sometimes still have what the fuck moments with your AI use?
And if so, do you have an example?
Yeah, so obviously I have it a lot.
Es ist also interessant, dass die ganze Debatte über Vype-Coding ist, weil viele Leute haben einen Erfolg mit Vype-Coding bekommen und viele Leute haben es versucht.
Und ich habe meine erste Applikation für den ersten Mal, nur die Cloud Model zu benutzen.
Ich habe eine einzelne Reihe von Coden.
Es war sehr gut.
Und ich habe viel confidence, dass das ist die Zukunft.
Ich habe in eine Zeit, wo ich keine Expertise hatte.
Der erste war mehr als eine Big Data-Application, wo ich viel Expertise hatte.
Und dann habe ich eine heavy und complex Frontend-Projekt und ich habe komplett verloren.
Same Model, same Person behind the Model und zwei completely different Outcomes.
Und dann habe ich angefangen, warum das ist passiert.
Und dann habe ich es.
Even in the first case, there were little mistakes, right?
So the model tried to trick me into certain situations, which didn't really make sense from the technical architecture point of view.
But because I was an expert there, I could immediately put my finger on the problem and I could say, hey, that's wrong.
And the model was saying, yeah, sure, okay, let's fix it.
Like, no worries.
And it tried to fix it.
And it said, like, now I fixed it like this.
And I said, like, it's still wrong.
And then with three or four iterations, we got where we wanted to go.
But when I went into a territory where I had no expertise in the frontend, I just relied completely on the model.
And probably the same thing happened.
It was something wrong there, but I couldn't put my finger onto the issue.
And that's the moment where you understand you still need the human, you still need the expertise.
So the model essentially is just an extension of your current capabilities.
And it's a lever which helps you to be more productive, but you still need to have some sort of an understanding of the underlying technology.
So that's one.
And obviously I had all the other examples which many other people had where the model tries to rewrite the test to fit a wrong behaving code.
Or the other thing where I tried spec-driven development and I created a huge spec in several hundred tasks.
And I said, okay, now let's implement it.
And then the model did like 30 and it stopped and said, okay, I will implement the rest of it when it will be needed.
And I was like, yeah, that's the reason why.
All of them are needed.
Nice story.
You are learning from humans and from human data and you have this deflection and healthy laziness in the models as well.
Yes, thanks a lot.
I can relate to this a lot actually.
Yeah, all right.
Thank you.
Yeah, I think we need to finish here, right?
Thanks a lot.
Super packed episode.
Thanks so much.
Vielen Dank.
Ich denke, dass keine Schleifte, die du irgendwo put hast, ist ohne Schleifte.
Aber wenn du genug Schleifte hast, dann die Chance, dass ein Risiko durch alle Schleifte ist, ist sehr, sehr hoch.
Und das ist genau das erinnert.
So, sie benutzen eine mixturee von verschiedenen Dimensions, die sie auf dem Fly generieren, Sentiment Changes, für zum Beispiel, und dann rune Statistik-Analysis.
Und diese Analyse...
After the fact is then being fed into further improvements of the agents.
Really, and that was just one example that they're doing a ton of things, quite impressive and really insightful.
What's one of your takes?
What stuck with me is actually.
The use of Mighty Agents absolutely makes sense to have some kind of task breakdowns because that would lead to smaller problems, smaller context, less context degradation and so on and so on.
But it was a nice reminder that you should not shoot a large project on the LM, but think about how can you break that down and then merge it at the end.
Absolutely.
Another thing that I also found quite noteworthy is the Prototyping for Requirements Engineering, what Shava described.
That's something that Anthropic also describes.
Fiona Fong in her talk in the recent Cloud with Code conference described something very similar.
And it makes so much sense.
And prototyping is very cheap now.
Hence, having it in the beginning before you commit any other capital reduces waste.
Absolutely makes sense.
Actually, it also doesn't matter whether it's clothed or lovable or bored or one, but really having a prototype rather than just thinking in theory how things could be absolutely a good point.
For me, it was also interesting to hear that they do not train their own models.
They're relying on probably one of the larger frontier models.
All they are doing is focusing on the harness, which is, I think, also a nice learning that you do not need to shoot with large guns on problems, right?
So a standard LLM is capable on a lot of things.
You just need to focus on the harness.
And they for sure do a lot when it comes to that and also all the fine tuning stuff.
But that was a good learning for me.
Good takeaway.
And it makes a ton of sense since they don't need...
Das ist ein vieles Grund, usually.
So, sie müssen sehr, sehr schnell transcribe das Wort in Text und dann act auf es und dann text-to-speech back zu den Kunden.
Absolut.
Und es also reduziert die Komplexität.
Denn in der Endeffekt es auch technologisch ist, und die mehr Technologie du hast, die größer die Komplexität ist.
So, es ist wirklich toll, oder zu hören, von solchen Leuten.
complex scenario they are in, that they rely on standards.
Yeah, that actually is a great segue to my third point that I find very noteworthy.
It's how they test their agents.
What he described is that they are simulating sessions repeatedly and then they're running evals.
Und das ist eigentlich wie du es mit LLM zu bauen.
So es sind die Produkte, die LLM nutzen.
Sie brauchen eVals und sie haben eine Infrastruktur für die eVals zu machen.
Das wäre vielleicht auch ein guter Topik für einen unserer Upcomingen episodes.
Es muss sich ein Fokus aufwachen, wenn du AI in deinem Produkten adaptieren, nicht nur für die internen Prozesse.
Für mich ist eigentlich nicht ein realer Insight, aber es ist toll, weil ich nicht weiß, dass Paloa ist mit der Mag7 in der Tatend zu haben.
As Schabba sagt, sie haben eine hohe Talent-Densität, die die 1% in der Industrie sind.
vielleicht eigentlich ein guter Ort zu arbeiten.
The Beyond Vibe Coding Podcast ist ein Projekt von Sebastian Heidemeyer zu Erpen und André Neubauer, Partnerin mit ImpalaSearch.
Der Content ist von uns und unseren Gästen.
Join the discussion on LinkedIn oder visit our website, where we publish all episodes.
Für Fragen und Fragen, feel free to reach out via LinkedIn.
Danke für deine Zeit und sehen Sie in der nächsten Episode.
