# Mastering Agentic Engineering And AI Workflows

**Podcast:** HMZE
**Published:** 2026-08-06

## Transcript

If you have just one unified system anyway, then you don't need any of these tools and you can be more productive.
So that's a theory.
And in two months, I might come to you guys and say, great news.
Please do that.
Or I'd be saying, you know what, forget it.
I was full of shit.
This didn't work.
I tried.
It didn't work.
Welcome to another episode of our new season of Beyond Vibe Coding, partnering with Impala Search.
The go-to tech and executive search agency in Germany.
In diesem Podcast, wir exploreen die transformational change in Software Engineering und Knowledge-Work in generell.
Ich bin Sebastian Heidemeyer, Zerbem, CTO at North.io.
Und ich bin Andrei, CTPO at Trusted Shops.
Great to have you back.
After talking about Local AI last time, today the focus is again on Mastering Agentic Engineering.
Correct.
Wir sehen uns beim nächsten Mal.
Beto, wir sind super froh, zu haben Sie.
Es ist sehr interessant, was wir in der pre-discussion diskutieren, dass Sie eine tech-companie, die nicht nur digital ist, sondern auch eine physische Föhrung und hat die Prozesse rund um das.
Wir werden das vielleicht ein bisschen später in der Show reden.
Aber bevor wir beginnen mit unseren übrigen Sektions, bitte introduce yourselves briefly zu unseren Listenern.
Ja, happy to be here.
Thanks for having me today.
I'm Pedro.
I'm VP of Engineering at Termondo.
I've been in the Berlin tech scene for over 10 years now.
I'm originally from Portugal.
I live here with my wife and two daughters.
And I worked for a few scale-ups.
You might be familiar with some of them.
I was head of engineering at HelloFresh before the IPO.
I worked at Contentful, a pretty well-known API for CMS.
Und ich war auch Direktor von Engineering an Geteer.
Ich habe ein bisschen Erfahrung mit der Skalips, die ich weltweit arbeite.
Und ich habe auch für ein paar kleine Unternehmen, eine Start-up, die vielleicht nicht so weit sind.
Ich habe viel Erfahrung mit der Small und der Large, und ich versuche, dass das Erfahrung zu bringen.
Unsere ambitionen sind, dass die Home und Beyond des Homemens.
Wir sind jetzt Amongst the top providers in Germany.
And the ambition obviously is to go European and to expand our footprint and to decarbonize more and more.
I think we're in the middle of a heat wave.
I think everybody can see how important our work is.
Yeah.
Yeah, that's me in a nutshell.
Awesome.
Thanks a lot.
Yeah.
And the good thing is also you're mainly, I think, in the business of heat pumps, right?
And heat pumps also help with cooling during heat waves.
So as a double benefit.
Ja, genau.
Ja, genau.
Ja, awesome.
Thanks a lot.
Und wie usual, wir wollen auch dieses Episode mit der Status Quo.
So wie Sie arbeiten?
Wie Sie arbeiten?
Wie Sie arbeiten?
Wie Sie arbeiten?
Aber auch in Ihre Managerial-Agearbeit?
Ja, well, als wir sprechen, da ist ein Cloud Code Session, das Sie an uns an all meinen Slack, Confluence, E-mail und ein paar mehr Daten sources.
Und als ich fertig bin, ich habe eine Briefing.
Es gibt auch ein Briefen über die Dinge, die ich möchte mehr über über die Dinge über die Dinge über die Unternehmen und die Unternehmen.
Das ist ein Passive-Way-Ton für mich.
Es hilft mir meine Arbeit.
Die Prioritization part, ich habe nicht in einem sehr lange Zeit.
Ich habe ein paar Wochen gelesen, und ich bin surprised, wie es ist.
Turns out, ein paar Jahre ist nicht rocket science.
Und es hilft mich zu halten, es in mind.
Ich benutze es für Coding.
Ich habe nicht vieles in Produktion in Termondo.
Ich finde, wenn ich mich selbst beginne, meine Gedanken auf das, dann ich mich einfach auf das.
Und ich vergesse es, wie ich mich auf das, wie ich mich auf das, wie ich mich auf das, und ich vergesse es für einen Moment.
Und ich also kann es in den Weg bringen und ich kann es ausprobieren.
Ich versuche nicht zu schipten Produktion Code zu arbeiten, aber ich benutze es ein viel zu überlegen, um die technischen Probleme zu verstehen, um die wir uns zu sehen, um mehr Grounding zu meinen Opinions zu sprechen.
Wenn ich mit meinen Heads of Engineering und meinen Staffen, dann hilft es zu haben, wir alle die gleichen Dinge, meine Management, Leier und mich.
Ich benutze es für Forschung, wenn ich etwas in der Industrie oder ein paar Änderungen in der Industrie oder ein paar Änderungen passiert.
Es ist wirklich einfach so, dass ich ein paar Änderungen aus senden, und dann fünf Minuten später haben, ein nice documenter, dass du kannst.
Und ich habe ein paar Änderungen, dass ich kodieren und das ist, wie ich kodieren kann, wie ich kodieren.
Und sie machen etwas mehr Proactive Dinge, wie in meinem E-mail Inbox und sie machen Drafting Messages und das Kindes.
Und lastly, ich habe eine massive Knowledge Base.
Ihr habt wahrscheinlich gehört, das LLM Wiki Konzept, das kam ein paar Monate ago.
So es ist basically das.
Ich war das für ein paar Jahre, und wenn Karpathien kam mit dem, es hat mich verbessert, mein Format, was nice.
Es ist einfach ein Massive Obsidian Vault.
which is structured and indexed a certain way.
And everything that I do, every conversation, every brainstorm, every project, it just gets processed and stored there for me to have available.
It really helps with interviews, with planning, with all these things.
So those are the main ways in which I personally use AI every day.
I think the risk sometimes is that you can spend too much time just chatting with Claude instead of making progress.
That's the trap.
But when you can keep yourself focused, it's just really good.
I feel like my awareness is much bigger than it was before that.
Ja, awesome.
Thanks a lot.
I'm also constantly working on my knowledge base, so I try to utilize something like an LLM Wiki approach.
Now the latest hype is OKF by Google, right?
Also wanted to look into this, but I still, on the other hand, have only recently managed to set up Obsidian, working on my, I use it on a...
CloudBot, like on the host, right?
And have a headless Obsidian there with a self-hosted sync plugin and then have it sync with my computers.
The app on the smartphone is not yet working, but I'm getting there.
So it's also a lot of time that I need to spend there.
But anyway, it already helps me also in my day-to-day work, definitely.
Are you using AI mainly for oder ist es auch experimentierend?
Was würdest du sagen?
Me, persönlich?
Both.
Ich habe ich, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, ich habe, Und diese Agents sind alle custom-bilden, weil ich wirklich, wie jeder Schritt der Agenten-Loop funktioniert.
Und jederzeit ein neues Modell kommt, ich einfach das Modell in die Agents und ich sehe, wie sie mit dem Modell behave.
Ich sehe, wie sie mit dem Modell behave.
Ich sehe, wie sie mit verschiedenen System-Prompten und Instructions- und Tool-Calling.
Und das wirklich hilft mir zu haben, was für was möglich ist und was hype.
Und ich denke, es ist eine Coincidence, dass ich das seit November bin, weil es Ende der Woche ist, dass das Waffe von LLMs kam, die sich aus, aber nicht mehr auszulteren, sondern besser auszulteren.
Die Gemini 3s, die Opus 4.5s, das ist, wenn du diese lang-rumpen-rungigen Agente hast, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die, die Ich habe das immer ein sehr viel, und ich habe das immer wieder.
Und ich habe diese Angelegenheit und ich habe neue Angelegenheiten, um neue Angelegenheiten zu versuchen, verschiedene SDKs und Tools, und auch Programming-Langrages.
Ich finde, das wirklich hilft, meine Meinung zu haben mit meinem AI-Team und auch mit meinen C-Levels zu engagieren, in der ich mit meinen C-Levels bin, in der ich mit dem, was möglich, aber auch mit dem, was möglich ist, was möglich, was möglich, was möglich, was möglich, was möglich, was möglich, was möglich, was möglich.
So try and navigate that from more information that's more grounded.
So you basically have a small eval suite basically to evaluate the new capabilities of new analysts.
Yeah, I think you're being generous by calling it an eval suite.
That implies way more structure than it has.
It's more like a crazy scientist's desk with all kinds of bits and pops on top of it.
But yeah, I use it to evaluate how they work, yes.
I also do eval loops, of course.
Aber ich habe ja gerade gesehen, ich habe einen Artikel von Steve Yege, oder ein Buch, wie er diese longen Buchstaben, aber immer sehr, sehr interessant.
Und der letzte Buchstabe ist der Flat Curve Society, wo er auch macht die Pointe.
Certain people didn't see a difference between Opus 4.1 and Fable 5, while he saw a massive difference.
And the reason was he did have these projects in his back pocket that he could pull out that other recent or older LLMs failed with, but Fable 5.
Das ist wie er die neuen Generationen von LLMs testet.
Und ich muss sagen, dass ich auch das auch was, so ich benutze es wie ich die anderen Modellen vorher hatte und ich nicht sehen, dass es eine große Unterschiede ist.
Aber so die Approach, wie er schon erwähnt, und auch andere Content, ist wirklich, wirklich gut.
Ja, das resonates.
Sorry, go ahead.
Sorry, Pedro, go ahead.
Ich sage, das ist einfach warum in November mein AI experimentationslose auch auf, weil ich hatte dieses Use Case in meinem Backpocket, das nicht nur so gut funktioniert.
Und dann, wenn die neuen Models kamen, das plötzlich wurde möglich.
Und dann, die Sache ich habe seitdem, ist eigentlich sehr einfach.
Ich finde, das ganze Konzept der Prompt und Skills ist falsch, in einem wahr.
Because it's very one-shot.
It's like, hey, can this agent do this one job?
And can this agent, you know, kind of loop until it gets there?
But it's very limiting because it's very, the way people do it tends to be really prescriptive, right?
It's these like really specific, do this, don't do that.
When this happens, do that other thing.
And what I was trying to see is, can I give an agent more like a job description?
Like, hey, your job is to onboard a customer and make sure that they, you know, Das ist ein, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, was ich an, Especially since Opus 4.5 came out.
Suddenly that started looking like it worked.
And then it started working better and better.
So I have these experimental loops at home where I have like a fictional SaaS tool that I built just to be like the guinea pig.
And I will plug the like webhooks from that thing to another piece of code that I wrote.
And just events just go in there.
And it just reacts to things and takes actions.
And it kind of looks like it's just a person doing a job, which is pretty crazy.
Obviously, the job has to be narrow.
And it doesn't get it right all the time.
But you know who else doesn't get it right all the time?
Us humans.
So that's why I saw that big leap.
And I can't say that for that use case, I saw such a big leap when Fable came out.
I think Fable is better at other things.
Aber ja, das ist ja.
Ich denke, dass es mehr von diesem Thema kommt, ich glaube, bevor das Jahr ist über.
Das ist mein Prediktion, aber wir können darüber reden.
To be honest, das Setup, das du mentioned, ist ein sehr sophisticated Eval Setup.
Ja, es gibt die typischen Evals, die man hat eine clear categorisation von Answers, und dann wird es automatisch gewartet, ob es ein Pass oder nicht.
Aber für was you are utilizing LLMs for.
This is probably the much better Evel compared to a standardized test.
So it sounds pretty sophisticated already.
But I think in the interest of time, we should move over to the main segment of our episode, right?
The meat, which is talking about specifically how you utilize, obviously, agentic engineering at the Mondo.
And then also, Ja, wie du persönlich verurscht die AIs, die sich mit verschiedenen Funktionen zusammenbringt.
Aber vielleicht starten mit der Art, wie du benutzt Agenten Engineering.
Ja, ich bin ja, ich bin ja, das ist ziemlich standard.
Wir haben, du hast, es hat sich angefangen, mit der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der Art von der TokenMax, aber wir haben die größten Spender in der ersten Zeit.
Das war einfach ein Weg zu drive adoption.
Es war immer noch nicht so unsustainable.
Aber ja, alle benutzen Cloud Code.
Man kann auch OpenCode mit verschiedenen Modeln aus der API auswählen.
Wir haben auch Leute, die eine Secondary Tool haben, wenn sie wollen.
Viele Leute haben Code und dann Codex.
Individuals use this in different ways.
We don't mandate a workflow.
We obviously have our software development lifecycle that everybody invites you by, but you can use AI in your own way.
There's no mandates to use it one way or another.
But we do have almost a religious schedule of knowledge sharing sessions.
Every week, unless there's something going on in the company, like a lot of vacations or some big announcement or something, unless something like that happens, every week we do a knowledge sharing session and two people come on, they demo or they present and they talk about what they discovered, what they're doing.
Und das ist ein wirklich guter Weg zu share, um, wie wir arbeiten.
Und es hat sich auch immer wieder getrennt, mit PMs und Leuten von verschiedenen Unternehmen, die die Unternehmen einfach nur zu kommen und sehen, was die PMs sind.
Um, es hat sich auch immer mehr getrennt, als ich erwartet.
Es gibt viele Leute, die jeden Tag, jeden Tag, jeden Tag, jeden Tag.
Wir machen Videos, wir sharen sie später.
Wir also...
Ich denke, dass alle anderen.
Wir haben unsere eigene, wir sharenieren.
So, das sind sehr standarde Worte in welchen wir, wie wir AI benutzen.
Eine wirklich gute Benefit von dem das mit dem Design Team ist, dass die Design Team, die sich able, quasi, aus Figma aus dem Figma aus, und mehr, aus Components.
They just make available, they have skills which know how to use those components.
And so an engineer can very easily, or a PM, can very easily put together a UI that looks and feels much like what our designers would have produced anyway.
And then when that prototype is ready to go to, or looks like it's actually going to production, then the designers can come in and make it a bit more intentional and everything.
But it really accelerates that design-to-engineering loop.
PMs are going nuts with AI, so now they can bring their ideas to life, they can make prototypes.
Prototypes actually work, which is a great thing.
They can actually take those prototypes to users, have them use the thing, collect information and feedback, and then...
Bring it to engineering to build in a way that's more secure and production ready and scalable and all those good things.
And we try to empower the PMs with as much as possible, like skills that tell them or don't tell them, tell their agents how to build it right from the get go so that the prototype gets as close to production as possible.
Das schaue, ich denke, es sind Leute, nicht nur die Termondo, aber die Idee, dass ein PM kann mein Job jetzt, ein Designer kann mein Job jetzt.
Was ich möchte, dass das ist, das ist großartig, wir sollten mehr machen.
Wir müssen mehr Menschen tun, weil wenn ein Person, die nicht ein Programmierer, kann man eine Produktion mit AI, was kann man?
Man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann man, man kann I think is always assured.
So let's actually focus on empowering as many people as possible.
The more software there is, the more there is for us to do as well.
So these would be like some ways in which we use AI at Termondo as a tech work.
That's pretty interesting.
And I would like to double click on some of the points that you mentioned.
Number one.
PMs are going nuts with AI and building prototypes and then they're bringing it to users and collect feedback.
How could you give an example?
Because I assume that it's not like a productive use, but rather something they use in order to see how their users react.
So it's more like an experiment setup or how are you doing it?
Es ist wie ein Experimental-Stereoids, weil du nicht nur bringen wireframes zu users und haben sie klicken und sagen dir, was sie denken.
Du bist übrigens in den Städten, die in den Städten sind, in den Städten sind, und diese Dinge haben reale side-effects.
So, jemand ist tatsächlich ein Wichtiges-Einigungs-Pieße mit einem AI-Prototype, Today, because now we have this kind of platform underneath everything that you can just plug into, like it's an API, right?
We didn't have this a year ago, and now we have this Termondo API that makes it just really easy to build whatever software you want on top.
And it kind of imposes rules so that you can't just go, you know, and create changes that you don't want to create.
So that gives a degree of safety.
Aber es ist natürlich nicht 100%.
So diese Experiments sind noch kontrolliert und sie sind nur mit ein paar Selected-User, die vielleicht ein bisschen mehr technisch-inclined sind, so sie können das mehr安全 sein.
Obviously für uns, Sicherheit ist eine große Herausforderung, und wir wollen die Envelope nicht machen, keine Obvious Mistakes.
Ja, so die Safeguards, das wir bereits bereits Also, in anderen episodes mit Fintech companies, help a lot.
So, if you have a proper setup, and in your case, it's also super serious, because there's actual hardware being involved, right?
So, you need to have safeguards in order to not do serious damage, sorry.
Then you can experiment easily, right?
Then you pretty much use it, if I understand it correctly, to...
Ja, correct.
I think that using AI does not mean that you have to be shy and it also doesn't mean that you have to accept security compromises, but it does mean they have to move deliberately.
und mit vollem Warennis von was du und was passiert.
Ich denke, das ist das ganze Thema.
Wir haben eine Sache, die ich möchte sagen, dass wir sicherlich hatten, wenn jemand etwas mit AI gemacht hat, dass sie vielleicht nicht haben, und es nicht mehr so ist, aber wir nehmen es wirklich...
Wir nehmen diese Dinge zu heart und, um, wenn es nicht ein properer Security incident war, da war ein Situation, dass jemand etwas zu haben, dass sie nicht verwendet haben, und sie nicht.
Aber es war ein Risiko, dass sie sich nicht verwendet haben, wenn sie die Safeguards nicht verwendet haben.
Und wir nehmen diese als learning opportunities, so wir nicht, you know, blame Menschen.
Wir nur admit, dass wenn man etwas off-track ist, es weil die System nicht...
und wir einfach nehmen diese als learning opportunities.
Und wenn man jemand mit AI macht, mit der AI zu beherrschst, ist es besser, wenn man sich sicher, dass man sich sicher ist, weil man das eine Liste ist, weil man das eine Liste ist, und man kann es nicht so schwer, dass man sich das Liste nicht so schwer, und man kann es nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer, dass man das Liste nicht so schwer.
Das ist alles wir können, und wir können uns einfach lernen und lernen und lernen.
Und du hast die Begrüßung zu vermeiden, was du?
Ja, das ist super wichtig.
Ich habe einige sehr viele, sehr viele Leute in meinem Team, die uns alle ehrlich sind, das ist wichtig.
Ja, ich denke, auf der anderen Seite, ich denke, dass es ein bestimmtes Failure-Culture ist, ich denke, ist super-beneficiiert, diese Tage, weil es ein sehr viel verändert.
Was ich was wondering, wenn du deine AI-transition hast, ist, ist es ein properer Projekt in place, um zu transitionen die Unternehmen zu etwas new?
Weil es ist ziemlich massiv?
Oder ist es einfach folgendert die Flow?
So, hast du eine bestimmte Plan in mind, um zu ändern, um zu ändern, um für die Engineering zu ändern?
How do you approach that transition?
Yeah, I wish that we were just going with the flow.
I think it would be a lot less stressful.
Let's just see what happens and adapt to it.
There's definitely an intention.
And I think that our intention is going towards Gentic, of course.
As a team, we want to kind of save ourselves as much time as possible so that we can focus on the things that matter more.
That's the mindset that we go into this with.
Ich habe zu sagen, dass mein Probably not going to be one where everybody just vibe codes everything.
I think that engineers are absolutely necessary.
And so our target state is not one where we don't code anymore.
It's one where we trend closer and closer to wherever it is that our attention should really be.
Right?
And we go back and forth on this.
Yeah.
We go back and forth.
Sometimes you feel like you make huge progress and that software suddenly can become almost autonomous and then you pick up a different piece of work and it turns out AI is just not as good at that thing and you have to kind of rethink the whole approach.
So I think this adaptability right now is really important.
100%.
I couldn't agree more.
And a question or maybe two questions in one.
How would you say just roughly percentage wise how much code is currently?
by AI in your organization and what's the target state?
So it doesn't seem to be 100, but it seems to be like asymptotically nearing the 100 or something, right?
That really depends.
So unattended code production, I think is zero.
Well, not zero.
Unattended code production is, I don't have hard numbers for this, let's call it 5%.
Was es ist, wir haben einige Agents, die sich für Sentry-Alerts anschauen, und sie suchen für root-Causeen, bei basically diffing, aufgängst auf GitHub und finden, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es, was es dass man so wie ein paar Dinge, so wie ein paar Dinge, das eine Agenten braucht.
Aber Code, das geht durch AI, 80-90%?
Ich denke, die Unterschiede ist, dass es alles gut ist, an dieser Stelle.
Wir haben wir mit Code Reviews gewohnt.
Wir sind sicher, dass wir ein paar Taktiken, die man kann, um die Börden zu reduzieren.
Aber wir haben uns für die Zeit, wir wollen die Menschen, die sich zu sehen, bevor es zu produzieren.
Das ist nicht ein Dogmatic Position.
Es ist um Accountability, in meiner Perspektive.
In der Target-State, ich glaube, dass ein Engländer nicht alle Länge sie shipt.
Ich glaube, dass ein Engländer nicht alle Länge sie schicken.
Ich glaube, dass ein Engländer nicht alle Länge sie schicken.
Aber die Engländer ist die Länge.
für alle linee sie schiffen.
Und das ist, dass ich mich auf die Fokus auf meine attention focusseite.
Wir haben diese Diskussion also in der Firma, wie es das bedeutet, dass wir uns zu Code Reviewen machen?
Warum wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es tun?
Weil wir es I want a situation where a person in my team target state, a person in my team can have a team of agents just build all the code for them and they're still going to sleep at night.
They're still happy.
They still feel safe.
They still feel like their work is under control.
They know what it is and they know how to fix things if they break.
I think the way to get there is going to be a hell of a lot more tooling around code production and landing.
That's super interesting.
Have you been looking into tooling around code reviews specifically and how you could automate this process?
Yeah, for sure.
We have some experience obviously with these kinds of automations.
There's a bunch of tooling.
I don't think it's necessarily worth mentioning many.
Ich würde sagen, dass die mehr wichtigsten sind die die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die, die sich die Das kind of things.
One keeps you safe or makes it less likely that something splits that might bite you and the other one helps you keep the code maintainable.
I think that with AI one challenge is it knows how to make software that works because all the training data and data annotation that it does is about having software which passes deterministic tests.
Ich denke, dass dieses Training Process für Models wirklich eine Funktion zu rewarden, die auch wirklich schwierig für uns Menschen zu tun ist.
Aber ich denke, das ist warum, wenn du Klaue ein Codebase und es zu verbessern, es geht komplett bananas.
Und an einem Punkt, du hast viel zu viel Code, sehr large Funktionen.
Wie man kann gut machen ist, ist nicht in Models.
Und dann gut code matters, weil man, als Mensch, es will, dass man einfach nur die flow des Daten und wie Dinge sind, wenn man nicht interessiert, dass man nicht interessiert, sondern man interessiert.
Du bist egal, dass man die Design des Designen hat, weil man auch die LLM-Sage beherrscht, auch die LLM-Sage beherrscht.
So diese Dinge matter.
Und Automation around diese Linten und Checks und Evalufe matters, auch.
Interesting.
We've talked to people from other companies that would probably say that the maintainability of the code from a human perspective is less important for them than from an LLM perspective.
And hence they're perfectly fine with automating more and more of this code review process so that the quality bar for LLMs is raised basically.
Obviously you have security, you have performance checks, you have the pretty much deterministic checks that the code is actually doing what it's supposed to do, right?
But code quality then is more a matter of can the future AI still work with it, refactor it and everything and then they would Like there was an example from Palo, I think they have these different kinds of agents who are part of the review workflow and then they have an orchestrator LLM on top, which then pretty much collates all the information and does a judgment.
So it's the judgment.
And if it crosses a certain risk bar, it would then still call out for a human review, but only above a certain.
Is that something you could see also for Thramondo?
Or are you still more focused on, still we don't know where the future is, so maybe it stays the way it is, but are you more on the path, no, no, the human needs to ensure that the human quality bar is met?
Yeah, fun fact, Italy is one of my best friends, so we talk a lot about this.
I kind of thought you might bring it up.
Ich will das.
Ich 100% will das.
Hier ist es, warum ich nicht.
Wie gesagt, wir haben nicht das API ein Jahr.
Für uns ist es noch immer noch neu.
Wir sind in diesem Grunde der Grunde der neuen Plattform-fürst, API-fürst, technik-Iteration von Termondo.
Das bedeutet, dass es wirklich wichtig für uns, jetzt, zu beherrschte und wirklich vorsichtig über die Fondations- und wir legen.
a lot of human eyeballs on the code, a lot of human eyeballs on the interfaces, and how this whole thing is put together.
I think we're approaching the point where, you know, even though we kind of just launched it, I think that we're going to have enough maturity in terms of the system being productive and stable and people knowing how to use it enough, that I think then I'm happier to give it more slack and just, you know.
Then really go to a system where you have agents going around and prioritizing pull requests and assessing risk.
Because if people understand the tool well enough, we can recover if the agents make a misjudgment.
And I think that's true of any agentic development.
If you have a really well understood, well observed, well behaving system, it's way safer to give agents more autonomy.
Das ist nicht neu.
Ich meine, als ich an der Head of Engineering at HelloFresh war, wir haben die gleiche maneuver, wir haben eine API, dass alles in den ATIen war.
Und ich habe das, das ist 10 Jahre alt, richtig?
Wichtig vor AI.
So, ich habe das, dass ich eine Design- und Procedure habe.
People could just stamp out features, right?
They had to think only about the business logic, but everything else about how that business logic plugs into the overall system, you almost don't have to think about because the structure's all there.
And so your productivity just shoots up.
And I think it's the same thing now with AI.
So we're getting to that point, I think, where we can do that.
I can't wait to try it personally.
I saw what Ito did.
I love it.
I want to try it for sure.
Got it, got it.
That makes a lot of sense.
And I just want to say that this is exactly, I think, how every company needs to approach it.
Parloa, obviously, running agents in production.
So agents is their business.
So they are masters in building agents the right way that do what they should do.
So for them, it's much easier to.
Build these agents also for their internal purposes, also for code reviews, right?
So it makes a lot of sense.
I couldn't agree more.
And I think, Andre, you had one final question, right?
And then I think we need to slowly but surely come to the end.
I was just wondering if you just made the statement that agentic engineering requires certain prerequisites, if that is going in that direction.
So, if I got you right, you say, given the limitations in your platform, you haven't got far in the transition.
Is that correct or did I got you wrong?
I think it's directionally correct.
The point that I was making was that to...
To give agents more freedom, you need to have a more solid foundation.
I think when the foundation is not solid, or even if the foundation is solid, I do believe that ours is.
But the engineers don't have, I think, enough time at the wheel with the new platform so that they can actually understand it well enough so that they can better recover, like if an agent goes amok, for example.
So I think it's a mix of just the pure solidity of the underlying platform.
und der Menschlichkeit mit dem bestimmten Plattform.
Ich finde das sehr interessant.
Ich habe versucht, das zu formulieren, dass ich das richtig habe, weil so far mein hypothesis war, LLN ist nicht so viel, besonders die neuen Models und die neuen Generationen, sie nicht so viel über die Strukturen.
Qualities and so on and so on.
So need to think about that.
But actually also brings me to the point, Sebastian, we want to record a session on architecture and agentic engineering.
So what is needed to make that work?
Maybe we invite you for that session as well, Pedro.
I have opinions on that.
So yeah, I'd love to join.
Awesome.
Sounds great.
But I think you're fundamentally right.
I think it's a matter of do you have the right safeties around?
Do you have strong enough harnesses and tests and interfaces?
I think if you have those, then it almost doesn't matter what's going on inside the boundary, right?
I think the question is, what's the boundary that I want to pay attention to?
True.
For me, that boundary is still kind of slightly low down the stack.
But I'm trying to push it up, for sure.
Makes sense.
Awesome.
And with that, I think we should come to the final segments of our podcast.
One thing that we're always talking about is the reality check.
So what are the what the fuck moments or the wow moments that you still sometimes experience with LLMs?
I think the Es ist so easy für einen LLM zu werden, noch zu verändern.
Ich sehe das mit meinem Chef-A-Staat.
Jetzt funktioniert es sehr gut mit einem Agenden.
Wenn ich den gleichen Agenten sagen, sie sich einfach aus den verschiedenen Dingen verändern, dann wird es sehr schnell verändern.
Und dann ist das wirklich schlecht, dass sie ihre eigenen Worte als Fakt nehmen.
You're brainstorming something, they tell you something back, and if it's wrong, that session is poisoned.
Because at some point, they will cite themselves back.
So my chief of staff has just heavy, heavy use of subagents so that I can compartmentalize the context as much as possible.
A wow moment, though, is this.
One of the biggest challenges for me as an engineering leader is that things are scattered, right?
Like the same subject, same topic can be on Slack, on email, on Confluence, on GitHub.
You're going to have evidence and conversations that all pertain to the same topic happening all over.
And so what I did was, all right, I'm actually going to like periodically, you know, grab updates from all these sources, and then I'm going to have an LLM agent just...
pass through them, actually multiple passes, but basically coalesce these updates into either existing threads, I call them, or new threads.
And this enables me to actually have way higher situational awareness because some code might have landed on GitHub that relates to some discussion which happened on Slack that has something to do with an incident or like a higher error rate on Sentry.
LLMs are surprisingly good at piecing that together.
So that to me was a wow moment when I was like, oh, hang on a second.
I can actually have like, you know, Telegram updates coming into my phone when things really matter and things that don't matter so much.
I just check them out later, but they're all correlated.
That's pretty cool because it's harder to do that as a human, I think.
Absolutely.
Because of the sheer volume of information that's flowing in, right?
Pattern recognition capabilities of AIs, LLMs, that's what they're made for, right?
Perfect, thanks a lot.
And as a final point, do you also have a prediction for our listeners?
My prediction is not original, but I think that we are closer than we think to having, let's call them synthetic workers.
I think a lot of office work is very procedural.
This also includes a lot of what I do.
There's a lot of stuff that I do that honestly doesn't really need a human to do.
It's very procedural, very chore-like.
I think LLMs have been good enough to do a significant amount of this work since November last year.
What I think is missing is tooling and the ability to discriminate between different situations that look the same but are not.
Und dann, wenn wir das, das nicht so schwierig werden, werden wir werden, dass Leute, die wirklich haben, wie Arme der Regen, die sie kontrollieren, um alle kinds von Arbeit zu machen, haben wir über das ganze Jahr gesprochen.
Ich denke, wir sind viel zu viel zu haben, als wir das wissen.
Und die wichtig ist, dass...
A lot of what we do is not that complicated.
We just need better ways of giving agents the information and the tools that they need to act on it.
Like if you have to pick up the phone and call somebody to do something, obviously that's still difficult, even with voice models, even with all that stuff.
If you have to go through five different systems, one spreadsheet, Salesforce, AppSheets, you know, some whiteboard that somebody's keeping somewhere.
Das ist natürlich off-limits zu Agents in praktischem.
Aber ich denke, dass wenn man all die Daten zu der gleichen Surface macht, Agents kann man sich nicht mehr überwinden.
In fact, bei Termondo, wir leben in der Zukunft, in einer Art, das ist auch eine andere Grund, warum wir die neue API verplatformen haben.
Weil wir ein Jahr ago gesagt haben, wir können das Ganze sehen, dass wir alle unsere Daten zu werden, die wir alle Daten zu werden, die wir nicht mehr überwinden.
Und wir sind schon wieder automating.
Wir sind ein paar Arbeit und das macht es viel mehr produktiv.
Wir sind am Ende.
Ich muss jetzt fragen, wie Sie das?
Wie Sie machen all die Daten zu den LLMs?
Ist das bei MCPs?
Sie haben eine Central-Platform?
Oder müssen Sie wiederholen Sie?
Ich bin froh, aber es gibt MCPs, die mich die hellen aus, weil sie zu slowen sind.
So we're moving towards more like hybrid search knowledge bases.
But the way that those are usually done is a little bit naive.
And so the way that you should use those is way more agentic.
Like there should be an agent in the knowledge base that is like interpreting your ask.
Taking action to answer it, but that's a whole other topic.
But yeah, knowledge bases, hybrid search, and agents trying to decide how to get you the right information.
The trouble with this, permissions.
But once you can search your whole confluence in five seconds instead of a minute, that's pretty transformational.
Hybrid search is, again, not rocket science, so we actually have two or three different.
Das ist ein Problem, wir haben ein paar Unternehmen, aber wir haben ein paar Unternehmen in Deutschland, aber wir haben ein paar Unternehmen in Deutschland, aber wir haben auch ein paar Unternehmen in Englisch.
Und zu suchen über verschiedene Sprachen ist nicht schwer, aber es ist ein Problem, wir haben zu denken, wir haben es zu lösen.
So the new knowledge base that we're building is more unified, can work across all these different languages and domains.
And also like the agentic part of it is like people often ask multiple questions in the same query, you know, like what are the different models that we serve and what's the process to install each?
And this is going to live in different parts of the knowledge base, right?
So if you just do a naive query, you get nothing or you're getting complete information.
So the agent has to...
Take your question, actually check if it's multiple questions or just one, ask in different languages, call us all the answers, all that kind of stuff.
But what I think is really interesting, and I did an experiment about this during my vacation because I can't stop.
And I actually am trying to build this as a prototype internally is...
When you actually start putting hard business data in this knowledge base, right?
So the knowledge base is not just pros.
It's also like tabular data.
And you give the agent that knowledge and you give it this dual mode of operation where it can either or both try to get narrative answers to questions or just retrieve data.
When it can do both things and it's all in the same place, so it's fast, it is kind of crazy.
Wie kann man schnell, und bei schnell, ich meine unter fünf Sekunden, eine sehr komplette Antwort zu, was es ist.
Und wenn man das so macht, die andere Frage, die du erst fragen kannst, und das ist, was ich wirklich möchte, ist, wenn ich all diese Daten in der Systeme, und es ist Strukturiert in eine Art, wie viel ich noch brauchen, die anderen Systeme, die Daten aus dem Systeme, Was ist das in die Systeme?
Das ist so viel, was ich habe, um, in einem Jahr zu payenden, um, ein Jahr oder so?
Und die Antwort ist, viele, natürlich.
Da sind viele, natürlich.
Ich bin nicht sagen, SaaS ist es.
Aber ich bin sagen, ich denke, das ist ein ganzes neue, um, das nicht genug genug sind, und das könnte eigentlich ganz transformational sein.
Weil das die Nummer-one problem ist, ist informationen, die sich nicht mehr aufbauen.
Und du hast entire industries, wie Looker, Bigquery, die hier zu überlegen, das Problem.
Aber ich finde, ich wirklich mag diese approach.
Es resonates mich sehr mit mir.
Ich sage immer meine Leute in meinem Unternehmen, The systems of record don't start working on something with having in mind that you want to get rid of the system of record, but rather making the information more accessible to agents and then more easily accessible to everyone.
If in the long run you can replace certain systems of record, I would say that probably analytics databases that hold tabular data have a purpose because what you...
What I have in a knowledge base is probably the collated information from different data points that you don't need always and every time.
But still, if you do the first thing, then you can contemplate about the second thing later on per system of record.
So it makes a lot of sense.
One question, what qualifies as knowledge in your definition?
Oh, that's super fuzzy.
I'm struggling with the same question.
I have an answer.
You have an answer?
Okay.
Actually, I was in knowledge management when I was a working student back then.
And there, the categorization was as follows.
So you have data, which is individual data points, basically.
Then you have information, which is derived from individual data points, right?
Layer would be knowledge, which is pretty much, if you wish, derived knowledge from different pieces of information.
So that's the categorization in knowledge management.
I love the definition, right?
I often work with that definition as well.
I think the reason why I'm struggling is because I think we're calling these things knowledge bases so that we don't call them databases because that term is already taken.
Aber ich denke, dass die Art der Knowledge Base, die ich denke, includes all die drei drei Fälle.
Und es ist die Aufgabe der Agenten zu versuchen, was man an welchen Punkt ist, und das ist wo die Komplexität lives.
Und es ist auch wo die Usefulnesse kann, weil, für example, die Sache ich wirklich auf jetzt arbeiten, hat viele SOPs.
Standard operating procedures for customer support, for installation, for all those things.
So you might call it a knowledge base.
Like this is derived from other things and it's end user consumable.
But at the same time, I built this prototype months and months ago where this is like a, as I said, I built a lot of agents for fun and one of the agents, I actually built a whole platform around them and the platform gave the agent the ability to build its own database because it was working with a lot of data, right?
Like just hard data, just numbers and figures.
And at some point it was kind of stupid to keep asking the target system to give you all the data at all, like all the time.
So I just built a data store.
Once I built a data store, I remembered my Contentful experience where each customer gets their own custom data content model.
Und ich war, was wenn die Agenten könnte seine eigene Datenbüder?
Was wenn die Agenten könnte, hey, ich habe diese Enten mit diesen Ressenten und diese Daten zu tun?
Das ist super, super Trivial.
An Agenten kann das super gut.
Es ist ein sehr thoroughly studiertes Kindeskundesign.
Und so die LMS sind wirklich gut at das.
Und das war eigentlich ein wow-moment.
Das war ein wow-moment.
Ich war, hey, du bist ein...
Ich weiß nicht, was es.
I built an agent to try to find me a doctor because Berlin, some specialties are hard to book with.
And so this agent just was able to send emails and listen up for replies and keep track of who had replied and who hadn't.
And the agent just said, all right, I'm going to build myself a little database of doctors and messages and replies and statuses so that I can remember this outside of my context.
Because if you rely on context, it can rot.
So the agent just saying, hey, I'm going to build a database and I'm going to go off and do this until you're either booked or on a waiting list.
That was a wow moment for me.
And this kind of dynamic database, I looked at that and I thought, this is basically Salesforce.
A lot of the complexity of Salesforce is what is your data model and what are your data flows?
Und es ist, dass die Database einfach nur ein paar Text-Files, Md-Files oder JSON oder whatever, CSV oder...
Ja, die Workflow-Behaviour, die wirkliche Dinge sind.
Und wieder, ich sage nicht, dass ich eine Verlust-Fersource mit einem Vybe-Coded-Experiment habe, aber ich sage, dass diese Tools zu einem sehr einfachen Primitivismus sind.
Und die AI ist eigentlich sehr gut an creating und manipulieren diese Primitivismus.
Und ich denke, das ist wo die interessante Entwicklungen kommen.
Ja, für jeden Fall.
Ich denke, wir müssen jetzt zu Ende kommen.
Danke viel, Perdo.
Es war toll, mit Ihnen zu sprechen.
Und wir sehen uns nächste Mal.
Danke, guys.
Bye.
Join the discussion on LinkedIn or visit our website where we publish all episodes.
For questions and inquiries, feel free to reach out via LinkedIn.
Thank you for your time and see you in the next episode.
