# Shifting Bottlenecks Left in AI-Driven Engineering

**Podcast:** HMZE
**Published:** 2026-08-20

## Transcript

You pretty much jumped ahead, so it's not the review any longer.
You say basically the pipeline is so good that solid software is coming out, but the bottleneck actually has been revealed and being before that.
So you shifted the bottleneck left and this was the PRD quality, right?
Exactly.
Welcome to another episode of our new season of Beyond Vibe Coding, partnering with Impala Search.
The go-to tech and executive search agency in Germany.
In this podcast, we explore the transformational change in software engineering and knowledge work in general.
I'm Sebastian Heidemeier zu Erbem, CTO at Ecosia.
And I'm Andrei, CTPO at Trusted Shops.
Great to have you back.
Today we are talking again to a very senior tech leader.
We are talking to Jürgen Helmers, precisely Dr.
Jürgen Helmers.
SVP Engineering at Andercore, and many interesting past positions.
Yes, Jürgen also went through the Paloas School, like some of our guests already also have been from the Paloas team, and he's implementing loop-based engineering across the engineering org at Andercore.
He gives detailed explanations of their approach, and it's really interesting and worth a listen, so I hope you enjoy.
Welcome Jürgen Helmers to our nice podcast.
And we're happy to have you and would like to ask you to introduce yourself to our listeners like usual.
Sure.
Thanks Sebastian.
My name is Jürgen.
I'm VP Engineering at Enderco.
It's a small startup here in Berlin.
Originally, I'm a biochemist.
I have a PhD in biochemistry, have done a lot of work that is really computer intensive, scripting and building 3D models.
Ich bin der Berlin Startup-Scene als ein Ruby on Rails engineer.
Ich habe dann einen Manager, habe ich mehrere Stints an Unternehmen.
Ich habe Wafair gearbeitet, HalloFresh, Parloa, und EnderCo, und jetzt bin ich in EnderCo.
Das war ein Concise.
Vielen Dank.
Super cool, also, dein move aus der Akademie und dann in der Engineering.
Das macht also Sinn, weil du in der Chat hast du, als ein Scientist, wie viele Wissenschaftler sind, also das macht.
Aber das also dann auch, das du hast, dass du, ja, du hast du, also, du hast du, also, du hast du, also, du hast du, also, du hast du, also, du hast du, also, du hast du, also, du hast du, oder?
Ja, du hast du, also, der erste, der ist, der, So how do you work?
How do you use AI?
What's your tech stack personally?
Yeah, that's an interesting question.
For the longest time, definitely at Wayfair and at HelloFresh, I was a people manager and building strategy and kind of like thinking about what can we do beyond the quota.
Endercore ist sehr unterschiedlich, weil es eine sehr unterschiedliche Firma ist.
Wir hatten Teams für alles.
Da war der SIE-Team, da war die Platform-Teams.
Die distribution der Verantwortung war sehr weit und da waren verschiedene Teams mit einem sehr unterschiedlichen Purpose in der Firma.
Hier bei Endercore ist es ein klassischer Start-up.
Du wehre die Hände und du machst das eigentlich für die ganze Firma zu erreichen.
So, when I started Endercore, we had the need for increased velocity in engineering.
And so, it would have been one thing to make plans of what it actually is we need to build.
And I also did and still do that.
But at the same time, I actually realized that with AI, I'm empowered again to actually contribute to engineering.
What we're using is the entire company is cloud-based.
So everyone at Enderco has a cloud account and there's instructions on day one on how to set up your co-work, your working style and your role definition in order to get you effectively using AI.
And there's a couple of tips and tricks that are part of your setup document that actually encourage you to not spend time on redundant tasks.
Every report I have to file, the weekly update that I give, the evening one-minute email that I do write, everything is based on an agent.
I do a lot of recruiting because we're hiring at Enderco.
And I write a lot of scorecards and they have a certain format.
So I created a skill that takes the transcript of every interview as well as my personal notes where I only indicate yellow or red flags.
And a scorecard is automatically created against our career framework.
I use it very often, daily almost.
I only had to override it twice, so it's actually working quite well.
And it's essentially the guidance at Enercore is if you do it the third time manually, create a skill.
Das macht viel Sinn und ist auch ein sehr guter Guadalign für die ganze Team.
Und wenn ich verstehe, das ist für alle, alle.
Es ist wirklich alle.
Alle, of course, die Idee ist, dass wir in Tech, weil wir AI sehr intensiv sind, also in der Software Development Lifecycle, dass wir ein bisschen von Experten, wenn es zu AI geht.
So, Prompting Techniken, Fuchshot, für example, das ist nicht wirklich ein paar Leute, so wenn jemand Probleme hat und ist, dass sie sich in der Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Rechte auf den Re Wie alt ist Endocore, und wie groß?
Endocore ist fünf Jahre alt und hat in den ersten April dieses Jahr.
Es ist sehr recent, dass es aus dem Jahr kam.
Wir haben uns gerade 100 Personen geholfen.
Got it.
But still at a decent scale, right?
It's not like another 10, 20-people company.
No, we don't fit into a small room.
We have multiple areas in the office, so it already has a certain size.
And it is also that people use AI, just to maybe extend a little bit on the topic that I introduced before of how we use AI.
Very recently, a sales agent actually created...
Das ist ein Tool, das Problem für ihn hat, das wir auf unserer Roadmap haben.
Und natürlich, es ist complex.
Es integriert sich in unsere Backhouse-Office-Tool, die Salesforce ist.
Es hat, wie die Transcripts sind, eine Qualität, und es macht es wirklich, wirklich kompliziert.
Es ist sehr, sehr gut, wenn man die Zeit, um es zu implementieren.
Es ist sehr, sehr gut.
Er hat eine Quick-and-Dirty-Solution gemacht.
Er hat die Skateboard geplant, während wir die Fahrer planen.
Ich habe es und habe eine Tricycle geplant, die ich heute aus dem anderen Sales Agenten habe.
Wir haben eine Hoffnung, dass das hier tatsächlich lösen unsere immediate Probleme.
Es ist ein klarer sign, dass nicht nur in Engineering, aber auch in Sales, die Leute denken über AI, sie versuchen es, und sie versuchen ihre eigenen Probleme mit dem, und sie experimentieren mit dem, was sehr schön zu sehen.
This is a great culture.
I need to ask for, like, for on own interest, like you're running Cloud Native, like is there any other platform you're running to make adoption easier for non-technical people?
We are using Cloud Desktop, obviously for everyone who doesn't like the command line, right?
So, but essentially it's the same thing.
When I joined in April, we still had Copilot, but we realized no one is really using it.
So we had Cursor, and then we realized no one is opening that except looking at Markdown files, but no one is actually actively working with it.
With the introduction of Cloud Code and the harness, it has been completely replaced.
Yeah, thanks.
Which, of course, when Fable was canceled, that's an indication of how dependent...
One actually makes oneself if one has all eggs in one basket.
Because if they should decide that a token now costs €70 instead of €7, a million that is, we would really struggle, of course, and might want to evaluate open source models instead of a Cloud model.
You can still use Cloud as an application and just plug in another model, but of course you would need to do some testing.
So currently we are dependent on Anthrop.
And your thoughts on this?
as you just laid out, to then quickly use an open source model, via open router or something?
I mean, at home, obviously I have been experimenting with OLAMA and I improved my digital personal document filing and the OCR aspect of that with AI by using OLAMA because I didn't want my bank statements and my personal communication to hit the cloud.
Es ist sehr schnell, aber das ist mostly wegen der limitationen der Maschinen, weil ich habe einen Mac Mini, also keine große Grafics-Card, das könnte eine mehr advanced Modell sein.
Ich denke, die Trend ist ein interessanter.
Die EU ist komplett dependent auf den USA, auf das eine.
Und das ist etwas, dass ich alle wissen, dass sie nicht mag.
Ich sehe keine großen Effizien.
To actually change it and invest into it, which I think you would need to.
The cost to develop these models is tremendous.
And I currently don't see anyone who's prepared to actually invest into it just to have an alternative.
I mean, to be fair, there recently have been some releases of models, actually even German models, right?
I wouldn't say that these are like models that can compete at the highest levels, not even at the not highest levels, right?
Aber ich bin nicht sicher, ob die Model Layer ist die one, die wäre relevant in terms of die Souveränität.
Als es gibt, dass es Open Source Models mit Open Weights gibt, dass man ja, muss man sich um die Garten zu bauen, um die Dinge zu machen, wenn es um die Models zu pre-trainer Daten gibt.
Probably not all answers about Taiwan or Tiananmen place will be correct, but at least you have models that you can use for agentic use cases, coding, whatever, right?
So this is probably possible.
And therefore, you basically just need the compute, right, in order to run.
You do?
I mean, ever since the 1st of April Source Code League?
Ich glaube, es ist ein Smartes Stück.
Und die Relyenz auf die Model ist, dass man sich eigentlich an den Modell verwendet und man kann andere Modell mit den Modellen verwendet werden.
Das ist etwas, was ich noch nicht mehr so machen.
Ich meine, ich bin der Wissenschaft, ich würde gerne einen anderen Setup haben, um zu testen und haben einen Matrix.
Es ist nicht so gut, aber was ist nicht so gut?
Was funktioniert?
Was funktioniert?
Was funktioniert?
Was funktioniert?
Wenn wir überhaupt nicht in der Situation kommen, dann ist es sicherlich möglich.
Ich glaube, es ist möglich.
Ich glaube, es ist wirklich möglich.
Ich glaube, es ist einfach nicht möglich.
Ich glaube, es ist ein bisschen schwierig.
Ich glaube, es ist ein bisschen schwierig.
Ich glaube, es ist ein bisschen schwierig.
Es ist ein bisschen schwierig.
Ich würde sagen, dass die Art und Weise, die man die benutzen, die man die benutzen, die man die benutzen, die man die benutzen, Specialized agents, the loops, the skills, it's probably just fine going for a GLM, for a Kimi K3 or whatever, right?
Also, I mean, at EnderCore, where we have an agentic layer that we actually use in production, there we don't use Cloud at all.
We, of course, make it work with larger models and then we ever so slowly walk back to actually the cheapest model that can actually do the task in good enough.
So we're actually using Flash Models, we use OpenAI, we use the Gemini Models, as small as we can make it.
Okay, but also there you have a certain dependency.
You just can choose.
Yes.
Okay.
I'm wondering if that might be.
Ich bin wondering, ob das vielleicht ein Spielwerk ist, um für Engineering zu lernen.
So, die andere Weise, nicht testen, nicht testen für Produktion, sondern die andere Weise, um etwas in Produktion zu nutzen, um dann zu nutzen, das Internationale zu nutzen.
Du getest, was ich sagen?
Ja, ich gete viele Leute, of course, von Unternehmen, die ...
GPUs in Europa.
You can run the OpenAI compatible open source models and 95% is good.
But my current motivation to test those 5% difference is not really that high.
If you have something that works, you prefer to continue.
Alright, let's move over to the main section of our podcast episode.
And this time, Es ist über Loop und Harness Engineering in einem distributeden Organisation, wie du es.
So, ja, bitte erzählen uns über das.
Ja, wenn ich EnderCore und EnderCore, das Team ist, als ein kleines Team, die waren, ja, wir sind AI, aber sie waren mostly Prompt Engineering und mit ein bisschen Kontext Engineering.
by pulling in information from different sources.
I then upskilled the team to harness engineering by creating a harness that would connect them to Jira, which we use for project management, to the Wiki for architectural documents, and ultimately starting to put up guardrails for engineering best practices that we want to work API first.
I created an engineering vision, which was a markdown document that could be ingested for planning.
I quickly arrived at the need for that already existed in my former companies where you would think, what is the smallest spec file that can hand to an agent to have a one-shot success?
That it implements it and I have almost nothing that comes up during a review of the code, that the code is good enough.
And I soon realized that mostly during AI-assisted reviews, There were always defects, always defects that were listed as blockers and as majors, as in you couldn't put it into production.
And one of our Nearshore engineers, I think it was Alexei, if I remember it correctly, he created a deep review skill within our harness.
that would actually follow the best practices for Java.
Our main stack is Java.
We're using Java 25 Spring Boot.
So that's what they're using as well.
And they are currently tasked to actually extract and migrate our logistics domain out of our monolith into a microservice platform.
And they wanted a good and solid code review because AI produces a lot of code and it's very hard to review all of that manually.
Instead of actually creating a loop there, they posted the code comments to GitHub as reviews.
I saw the chance to actually create our first loop.
I changed and took his very good review skill and I defined a loop.
I defined the loop as such as that I have one agent that actually does the review and loop engineering obviously, you know.
communicates with the file system.
So that one agent that does the code review creates an issue MD file.
Then another agent reads the issue MD file and tries to fix the majors and the blockers.
Then hands it back to the review skill in order to have another round of review.
And I defined the gate for this initial loop.
And that was my very first time that I actually tried loop engineering.
I defined the gate as in there's no blockers, there's no major defects.
And I defined a circuit breaker that this loop should run maximum three times to not endlessly burn tokens.
To this day, we never hit the circuit breaker, so that never triggered.
And we have a review skill within our harness that you would still need to manually trigger at that moment in time.
Das würde erzeugen in eine bessere Code-Qualität, das würde eine normaler Human Code-Review geben und das würde greatly erhöhen die chances der diese Software-Sophie eigentlich operating, wie es expected, in Produktion.
Once das in place war, ich es eine step weitergezogen habe.
Und was ich nicht wirklich mit der Harness hatte, war es diese Waterfall-Releases für complex projects.
Really long documents that would cover essentially a roadmap item.
And it would work until the very end to then spectacularly fail.
So I experimented then at that moment in time with, I asked the agent to create meaningful milestones.
Then once I had the milestones, I actually realized I could implement these milestones like we would in the past, where you would like milestone by milestone.
und start implementing these.
And then I did realize that with the review skill that I could put all of that in the loop, that I would have a milestone run and then run the review skill.
And if the review skill actually revealed the gates agreed, that I could then create a pull request in order for it to potentially start the next milestone.
I hooked it up to Jira.
I made a Jira base so that multiple people could actually work in parallel on a project by having an epic per effort, per loop.
And I moved all of the files that are necessary, the Claude MD file, the issue file, the loop MD file that captures where the loop is actually currently residing.
That all runs in a specific subdirectory so I can separate concerns and multiple people on the same repository can actually work on.
in order to parallelize.
I would move based on the loop's progress, the JIRA tickets, so I had observability in that point in my moment in time.
I would automatically update and create documentation, which is something that in my experience always usually fails because that's thrown out of the window first.
So it would know where our parent page for documentation actually is.
in order to ultimately create it if it doesn't exist or update it based on the progress of the loop.
I created this because I'm a big fan.
I have to admit, I've never done it myself very effectively because I was always a better coder than I was a tester.
I implemented this loop and the execution of the loop in test-driven development.
My friend David Sammert, he works for Equal Experts and he's a great coach for test-driven development.
Er hat sich die Qualität besser gemacht, so ich habe es mit der Qualität.
Ich habe eine Red Agenten, die erst die Testen, dann habe ich eine andere Agenten, die sich independently aufzunehmen, um die Red Testen zu machen.
Das ist die Green Agenten.
Dann ist es gehandelt auf die Review.
Und die Orchestrator dann creates eine andere Agenten, die sich die Fixing befinden.
So, alles läuft in ein separateer Kontext.
Ich habe ein Automated Stop in dem, wenn ein Kontext erreicht ist, ein Prozent, wo es nicht mehr Sinn macht.
So, vieles von mir in die Idee ist, um es zu entwickeln, dass wenn man weiß, was man macht, kann man eigentlich solidarzt werden.
Es ist sehr schnell und sehr schnell.
Aber es hat einen Achilles-Health, und das ist die Qualität der Plans, die man in die System sende.
It quickly revealed that our PRDs have a lot of gaps, have a lot of holes.
So it triggered some process changes in our company and our technical organization that we use Grill Me sessions to actually have an agent poke holes into our plans.
It makes them more consistent.
It generates a lot of open questions.
I realized that an agent produces questions in a way that stakeholders don't understand them.
So I built the skill that, depending on my input, it takes a PRD, a technical design document, and open questions.
And it assumes a role.
I can send it, hey, this is for the head of logistics, for sales, or procurement.
So it puts that hat on, and it avoids any technical language, and it starts speaking the language.
So that they have an easier time actually answering the questions that otherwise I would not get any answers to.
This way, we're shifting left.
We are in a state, and this is very new, so definitely for Q4, we want to make that our standard process is that we plan.
We plan PRDs.
Wir grillen sie, preferably zusammen.
Wir answeren die Fragen in der Sprache, dass unsere Stakeholder verstehen.
Und das ist natürlich mehr die Business-Questions.
Die Technik-Questions können sehr technisch sein.
In order to have a better understanding and a clear and sharp definition of what the scope is, what not the scope is, what the expectations are, what really needs to be implemented, which allows the loop engineering to do a better job in producing a piece of software that fulfills actually the stakeholder requirements.
Sounds impressive.
Yeah.
I think we need a moment to digest that.
Sorry for lecturing for so long.
No, no, no.
I think it was good.
Let's unpack it.
I think this is your point usually, Sebastian.
So usually when we talk about where the bottlenecks move because they move away from They usually move to review.
In your case, you pretty much jumped ahead, so it's not the review any longer.
You say basically the pipeline is so good that solid software is coming out, but the bottleneck actually has been revealed and being before that.
So you shifted the bottleneck left and this was the PRD quality, right?
Exactly.
And this is something...
I might refer to Dave again.
He has a paper route that I think is a very good one.
What is the effect of AI on the software development lifecycle on teams?
So it's reflecting on how the team behaves, how the team feels about it.
But also the fact that AI is greatly accelerating.
It's not only accelerating what is good about the software development, and that is you get faster to a finished product.
It is accelerating and identifying every single gap and everything that's actually wrong with your organization.
If your planning is no good, it's no longer revealed over the course of a quarter.
I have built an audit logging service with ingestion, with querying, and with SDK and Kafka support layer in 24 hours.
Because I went to the YOLO mode.
I experimented with that as well.
So it just went ahead.
I instructed it after every pull request creation, merge it and go to the next one.
Very risky.
A well-defined task, because audit logging, you find a lot of industry standards and good information online, which I included into my technical planning.
You can create things extremely fast.
So it's no longer a matter of months, quarters, or even weeks of sprints.
It's actually you reveal what's good and bad in a day.
And to the point is that we are moving through our roadmap faster.
We also doubled our velocity speed in the team, which also comes with exhaustion, I have to say.
And because it's continuous scope changes, I realized at some point I can have four terminal windows open at the same time and run four loops at the time.
With the fifth loop, I stop ingesting what I'm actually doing.
And I need to read and spend too much time in understanding what's actually going on.
And I can't answer the questions.
Fast enough anymore.
And I overwhelm myself.
That is the time where I realize I start dreaming about my loops.
That affects my family.
I have three kids.
And that's not a good thing.
So you need to slow yourself down.
It is one thing to be able to produce software fast and in decent quality, let's call it.
It's a completely different thing what it actually does to you as a human being.
Absolutely.
This actually reminds me of the Steve Yegge article, I think he called it the AI vampires or something, where he also suggested because of the higher degree of intelligence that our work now needs, so there's much more, like, less road tasks that we're executing, right?
It's much more...
Demanding from us and many more context switches were exhausting much, much faster.
He said that at 11 a.m.
he's pretty much done for the day and he can go to bed again.
And so he suggested actually a three hour day, work day maybe.
Maybe that's something for the future.
But then there's also other things at work that can be done and that should be done and that also add value that are less exhaustive, right?
Das ist ein sehr guter Punkt.
Ich war in der Lobby Area, an der Floor war, und ich war thinking, und ein meiner Peeren kam, er war, warum nicht arbeiten?
Er war, ich bin, ich bin, ich bin.
Und dann, um die richtige Balance in der Regelung, so dass die Business goals eigentlich met in Zeit sind, weil, am Ende der Ende der Ende der Ende der Ende der Ende der Ende der Ende des Engenheils, ist der Delivery und du hast dein Wohlbeammer.
Aber gleichzeitig, es gibt einen guten Design- und die System-Auswahl auf die System-Auswahl.
Denn die besser your design ist, die besser your plans sind.
Das ist ein Teil von mir.
Und du möchtest, dass du eine andere Teil in einem Microservice-Ensemble bist, das, am Ende der Tag, muss man zusammenarbeiten.
Und wie ist das eigentlich funktioniert?
Weil die Loop nicht mehr wissen, was es passiert.
Und das bringt mich zu meinem, hopefully, nächsten Interesse, dass ich noch mehr Zeit habe.
Wie machen wir es, dass alles wieder mehr mehr aware ist?
Wir bereits struggle mit sehr einfachen Facts.
Mein architectural vision hat keine Ahnung, was bereits existiert und was nicht.
So I actually need to update it.
So I started experimenting with automatically updating our vision to this already exists.
These changes have been made because otherwise it's very quickly outdated.
And then the instruction of what the big picture actually looks like is ever so slightly off.
And if an agent is ever so slightly off, it starts trying to fill gaps.
So I also started definitely during planning.
I became very, very strict and I included this section of.
You do not invent scope.
You need to ask questions to kind of like preempt what was actually happening without these kind of like guardrails that it filled gaps.
It put things together and invented formulas when it was coming to like implementation details where I needed to separate where does that actually come from?
And I couldn't find a source.
And that needs to be prevented.
It's that 15% of hallucination that actually diffuses and waters down every plant.
And one needs to be very careful there.
You need to be at high quality of the plant that goes in, but at the same time, you actually need to review the plants.
It's not enough to just convert it into a technical design document, then into a loop plant.
We need to shift left.
And this is definitely a learning out of, let's say, well, now almost two months of loop engineering.
You need to spend time reviewing your plans because just executing them is fast and easy.
But what if you built the wrong thing?
It consumes a lot of tokens.
It's very expensive.
And I got the goal from Philipp, our founder, that that is fine.
So there's a certain, of course, there is a certain limit, but we are not any close there yet.
But you really need to pay attention to what you're doing.
Otherwise, it's Vibe coding.
I can really relate to that.
like a solopreneur, I think that is easy also to implement.
What I find all the time very hard is to keep an organization in sync with that.
So what you just described is a lot of change, right?
And the speed at which this change happens is actually accelerating.
So I'm wondering how to...
Also, from an organisational perspective, how to ensure that the organisation can keep up with that pace.
Well, how do you do this, especially since Sebastian framed it in the beginning, you are also distributed?
Yes, we are distributed.
And we made the mistake, for example, of our current roadmap to focus on one domain because we wanted to extract it out of the monolith and move it to our microservice platform.
But that made it, with the advent of loop engineering, It accelerated so much and so many open questions piled up that we had exactly that one stakeholder as the expert who was suddenly faced with 100 open questions.
And that the engineering team that is distributed, they are working for an initial provider, they're actually located in the Ukraine.
They're very capable, they're extremely senior, which means they adopted this extremely fast.
They're actually contributing.
They're improving the harness, they're making it faster, they're making it more cost-effective.
But a side effect is that what I actually have them plan to implement until the end of the year, they're now ready next month.
Which means we need to get our A into G in order to actually come up with more plans so they actually have something to do.
It is accelerating everything.
It is a challenge not only for the engineering team to kind of like find the right balance between In the past, you would write manual code and that was relaxing.
It was something you enjoyed doing.
And it just came up in a retro last week.
That's like, this is missing.
Now it's just like everything is fast and it's extremely fast-paced.
And it trickles back to our entire organization that we need to think about faster in more detail what it actually is we all want and then map it back to our vision of where we actually want to go.
Which means we need to make faster decisions, not only in tech.
We need to make fast decisions in business with our stakeholders.
And they are sometimes overwhelmed by the speed in which we actually want to roll out.
Because suddenly we have like these three things that need to go live on the same day.
And it's like, no, you need to stagger it.
Because otherwise you don't have the time to monitor things.
And it's simply too fast.
I need to ask in a follow-up question on the technical thing.
Ist es eine Harness zu regeln?
Ich weiß nicht, wie ihr technische Foeprint an EnderCore ist, aber wie komplex ist das?
Und ihr habt dann eine sette, ich weiß nicht, eine Art von Skilz zu runen?
Es ist eine Kombination von Skilz, die eigentlich nichts but...
Markdown files, right?
And there's some connectors, they're MCP skills, or we actually have API credentials in a .n file that gives access to tooling.
It is a collection of skills that ultimately are interacting with each other.
So there is an agent that is orchestrating and fanning out.
So ultimately, it's a multi-agent setup, right?
So there's this agent that's always running the other skills and is delegating work.
In the loop recipe itself, I also implemented the use of different models for different tasks.
We use Fable for planning because that's where it comes.
Then we use the Opus models for the implementation as well as the review.
For anything that is interacting with third-party interfaces, whether it be Jira, whether it's GitHub to create pull requests, we're using the cheapest Haiku models.
Well thought through.
Thanks.
You mentioned that the speed has increased so much that it's almost going too fast and that it's hard to come up with enough plans to keep the team busy, which is like sort of a luxury situation, right?
Even though it like reveals another bottleneck.
Und du hast gesagt, dass die engineers in den Nearshoring-Teams sehr sehr senior sind, sehr experienced.
Das ist warum sie das sehr schnell verabschiedet haben.
Ist das für die ganze Organisation?
Oder soll man das bestimmte Dinge in order zu machen, um die ganze Team zu bleiben, und zu bleiben, sozusagen?
Ja, das ist eine sehr gute Frage.
Wir haben das schon mit dem, an meinem alten Unternehmen, wo wir gedacht haben, Wir zeigen, wie es funktioniert.
So in der Frontal-Knowledge-Sharing.
Und dann werden alle alle bemerkungen und werden adoptieren.
Das war absolut nicht der Fall.
Wir hatten sehr viele, die auf den Bandwaggen jumped.
Weil sie sie excited sind, und sie waren excited, um etwas anderes zu machen.
Altame ist es Verwaltung.
So haben Sie die klassischen Charaktere.
Sie haben die Naysayers.
Sie sind so, dass ich noch schneller bin, wenn ich meine Codee verwende.
Ich bin sehr gut.
And you have the yay-sayers, they are the ones that really invest.
But the majority of the group, the median, they can go either way.
And you really need to convince them.
So at Paloa, where I worked before, we went from frontal to actually a hands-on workshop.
I still remember Pedro actually sitting up front and guiding everyone through and actually taking everyone by the hand.
And I learned from that.
I adopted that.
So I started with knowledge sharing session.
I then took an example of, hey, this is how it actually works.
And I realized that people have very personal access to the AI.
They connect very different things with it.
While some actually do like to see them still as a craftsman that manually creates something.
And I was very much drawn to software engineering because I like woodworking.
Ich habe auch verschiedene Arten.
Ich habe die Gitarre.
So mein Ziel ist, dass ich meine eigene Spanisch-Gitarre habe.
Und hier haben wir einige, die wirklich die Manual-Aspakt haben und die Insistung auf es für ihre eigenen Vorteile und die Qualität, die man nur auf die Möglichkeit hat, man nur auf die Möglichkeit hat.
Mit der advent der Models-Generer und besserer-Aufgabe und uns zu entwickeln, eine Harness-Ausse, die ist letztendlich selbst-improving, wo man eigentlich Tell the hunters, hey, I noticed this and I also collect some metrics during these runs.
So I know where there's a lot of time being spent.
Just recently, I actually, I, well, I thought I reviewed and improved a contribution from our nearshore engineers that were supposed to save tokens and make it faster.
And I completely made it worse.
I burned token like ice in the, in, in the, in the desert.
Und es dauert forever eine Single Loop.
Ich habe also einen Menschen, der wirklich nimmt die Harness, die sich auf die Harness zu beherrücken und die Fokus auf, dass jeder pulleite ist, dass es eigentlich besser ist.
Currently, es ist ein Team-Effort, das sich in einem sicheren, ohne jemanden wirklich zu haben.
Es gibt viele, weil ich mich accountable bin, weil ich die VP of Engineering, ich sage, Jürgen, du musst es review.
Sometimes this was a pull request for like 94 files changed.
I didn't look at all of the 49 files.
I didn't have time for that.
So I was like, yeah, well, Alexei, he knows what he's doing.
I just fix it a little bit.
Turns out I was actually making it worse.
So in its sense, it is a system that by now you can actually ask it.
And this is the smartness, of course, of the models.
This is what I observed.
This is the expected outcome that I would like to have.
Improve yourself.
Und das funktioniert ziemlich gut.
Es ist sehr effektiv.
Und dann, zu gehen zurück zu den Knowledge-Sharing.
Knowledge ist mit der Harness, also die Menschen benutzen, die die Interface zu den Human nicht verändert werden, aber die Outcome ist besser.
Also, die Interface wird verändert, oder?
Ja, ich meine, die Interface wir nutzen, Sie sind in den Terminals.
Es ist eine Terminal-Application.
So die Interface ist alt-school.
Ich habe auch noch einen Teil davon, oder sie Menschen, die sich mit Precision beschäftigen.
Es ist so, oh, danke, und could du, bitte?
Ich meine, das ist ein Waste von Kontext, von meinem Perspektiv.
Ich meine, es ist sehr wichtig, wie wir als Team interactieren.
So das ist eine Herausforderung.
Es ist nicht nur eine Scope-Schalange, sondern es ist auch eine Herausforderung.
How do we actually communicate effectively?
And crisp and very clean instructions to an agent are extremely beneficial.
They're the death of every lunch meeting because that's not the way you communicate.
So that's another scope change that our team has to go through.
But using extremely precise language and be very specific what you expect and what you not expect.
Das ist das, was du eigentlich zu haben, die selbst-improvement und letztendlich haben die Results, die du wirklich nachher hast.
Und das ist das, was du musst, das ist das, was die Weise, die Harness itself ist, ist, ultimately, eine interessante network von verschiedenen Skills, die eine bestimmte Behaviour.
Und du könnte sagen, dass die Google search prompt, ich weiß, wie ich suche und findes sehr effektiv.
My dad, they don't, and they describe things in a way that leave a lot to the imagination, and they're very ambiguous, which means you're not finding what you're actually truly after.
And the agent is very similar, and the plans are very similar.
I, for example, noticed in one engineer, he used a plan to create an RFC.
Out of that, he wanted to create a technical design document, but he actually did a loop in itself.
He created an RFC based on an RFC.
because they didn't use very precise language.
That was a lesson for the entire team very early on.
It was only an hour that was wasted, but it was a good learning.
Ultimately, what worked for me, to go back to the beginning of the question, Sebastian, that you asked, is how do I get the team?
And this is why I got back into actually hands-on engineering myself.
I did realize I can talk about it as much as I want to.
I can share my, at that moment, to the team apparent theoretical knowledge.
Whereas they were actually asking me, it's all nice that you talk about it, but we don't know how to get started.
Can you help us to get started?
So I was sitting myself down with a hands-on workshop and then literally with pairing to support them to just get started.
Because like everything in life, the first steps are the hardest.
You need to overcome this kind of like hurdle before it becomes actually self-fulfilling, where they then realize is that Es ist sehr verbose.
Es ist vieles zu vieles zu produzieren.
Aber es ist eigentlich etwas, was es muss es tun.
Ist es uns dieses wow effect?
In der Zeit, wenn ich zu reviewen ein Stück von einem Staffel oder Principal Engineers habe, wie es, wow, diese paar Länge von Code tun all das, was es eigentlich supposed zu tun.
Und das war ein tolles Gefühl, zu sehen, dass es.
Ich habe das mit Code, das ist produziert von einem Agenten oder definitiv nicht von Alharnis.
Ist es eigentlich working?
Weil ich sicher es ist, weil ich habe diese Test-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven-Driven dass wir das wirklich was unsere Stakeholder wollen.
Wir wollen dann auch die Chance, dass das was aus dem Regressionen kommt, nicht nur können wir ein Button und wir runnen den Test suite, wir wissen, dass was wir es eigentlich beinwärtig ist, aber es ist das Wichtigste, dass wir das Wichtigste machen.
Wir sind ein Start-up, so wir brauchen Dinge schneller.
Wir sind nicht wirklich wirklich enterprise-software, wo wir schauen, wie viel Zeit wir brauchen, auf die Internetseite.
Es ist wichtig, dass die Features fast in dieser Phase des Codes sind.
All right.
Und die Frage, was du mit dem Kode zu sprechen und mit dem Team mit dem Team gesprochen hast, war das auch mein Impetus.
So über Christmas, ich habe viel Zeit mit den Agents in der Lage, wie ich es best benutze und wie es in unserer Domain und dann mit dem Team share.
Ja, ich meine, es ist ein Thema, aber wenn man dann kann es sich nicht selber machen kann, dann ist es ein bisschen wie ein Streetcrat.
Was war es anders?
Weil, du hast es, dass bei HalloFresh und bei Wayfair Leute gesagt haben, du musst dich nicht beutachten.
Was sind die Arbeit?
Du musst dich machen.
Hier ist es mein Job.
Du musst du etwas anderes machen.
Hier war es ein bisschen anders.
Ich habe es über das Bild, ich habe die Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bildung des Bild Das bedeutet, dass wir unser Leben schneller haben, wir haben mehr Zeit, um unsere Plans zu verändern, unsere Vision, was wir wollen.
Und sie haben uns das Speed gegeben, weil ich sie schon haben, besonders mit meinen drei Service, mit dem Audit Service Suite, ich soll sagen.
Audit Logging, ich muss es genauer sein.
Das war literally in 23 Stunden.
Und sie waren so, wow.
Das ist möglich.
Und kann man eigentlich nutzen?
Ich hatte einen integrieren test.
Ich habe diese drei drei Services geschlossen.
Und werdet ihr zusammen.
Ich habe es gesehen.
Ich habe es gesehen.
Ich habe es gesehen.
Das funktioniert.
Und jetzt wir starten mit dem Produkten.
Und es ist eine Sache zu sehen.
Aber wie ich gesagt habe, es braucht Hilfe zu werden.
Mit der Harness.
Weil du musst du es able zu.
Express yourself in a certain way in order to use it a fact.
Yes.
Amen.
All right.
I think the secret sauce to that, and maybe I just want to challenge that, also want to get challenged by you, Jürgen.
We often have these people in the podcast.
I'm wondering, what's the secret sauce to be able to step down?
Because this is actually what it is about, right?
So not people management, but getting back to software engineering, right?
And I'm wondering, is that the knowledge about software development lifecycle and all the techniques, I think in one of the recent episodes we talked with Ben Hoskins about that.
And he was also like...
My former manager.
Yeah, I know.
That's why I brought this up.
Wir haben so viel über die Techniken gesprochen.
Ich denke, wir sind mehr oder weniger in der gleichen Gruppe.
Wir wissen diese Techniken.
Ich denke, das ist vielleicht die secretäre Sache, warum wir diese Gruppe sind able zu stecken.
Ich bin wirklich gefragt, was die Grunde ist, warum die Managers sind, die Dinge zu tun, die Dinge zu machen, die sich in den Händen zu machen.
Do you have an opinion on that?
It's a very good question.
And yeah, I do.
It is ultimately when I used to be a Rubin Rails engineer, right?
So new versions coming out, I always thought, oh, let me just try that.
But just the setup of tooling, installing your gems and dependencies and getting something up, you could actually get into this flow of I can actually produce something was cumbersome.
Es war hart.
Jetzt, auf einem Agenten, die zu Kontext 7 ist, hat sich immer gesagt, hey, ich setze mich für Rubien Rails in dieser Version.
Es macht es.
Es macht es.
Es macht es und macht es das Threshold, dass es vorher war, es war hart zu jumpen.
Es braucht Zeit.
Du musst die Zeit zu machen.
Me having a family, having multiple hobbies, from playing instruments to going cycling, which is, André, you know that, you cycle yourself.
I sometimes disappear for these four or five hours to the dismay of my family.
I simply didn't have the time to actually fiddle then around with a computer and trying to set up the latest version on Ruby on Rails.
Now it's become awfully easy.
And I find myself on the sofa on Sundays.
Und ich kann mit es spielen.
Und ich liebe es, zu bauen.
Ich bin nicht ein Software Engineer und ein Computer Scientist, bei training zu werden, bei eigentlichememischemischemischemisch.
Ich liebe es, zu finden, wie Dinge funktionieren.
Und für mich, das Harness ist ein System, mit verschiedenen Elementen und verschiedenen Agenden, die zusammenarbeiten.
Und während meiner Zeit als Manager, es war mein Job, die die perfecte Team zu machen.
Und hier und da, ich habe es geschafft.
Ich habe die Team, die alle waren happy, ich habe sie gemacht.
Ich investiere in den Scrum Ceremonien und ich habe es geschafft, dass diese Gruppe von Leuten war in perfekt harmonisch.
Ich habe mich noch einmal, Runeisch, zu mir, Jürgen, ich glaube, wir können alles machen.
Und ich denke, das war, bis heute ist das schönste, was das Ding der Nisestes Sache.
Das hat mich als Manager gesagt, weil es all meine Investitionen in dieser Gruppe von Leuten eigentlich hat es geholfen.
Sie fühlt sich das alles, was was management was all about.
Jetzt ist es verändert.
Wie werden Sie als Junior, wenn Sie als Junior haben, wie Sie die Code auf Ihrer Seite haben?
Wie werden Sie diese longen Liste, wie werden Sie diese longen Liste von failuresen, die sich in die Technikale gemacht haben, als Sie als Techniker haben?
Das ist die neue Generation von Ingenieurs-Ingenieern.
Das ist die wichtigste Erfahrung, das Leben hat.
Das ist ein Fehler.
Das Hands-on-Interaktion mit etwas, das nicht funktioniert.
Du gehst über das Stack Overflow.
Du versuchst all sorts of wirklich old suggestions.
Und du wirst es irgendwie nicht mehr.
Das geht komplett weg.
Das ist die Angelegenheit.
Das ist die Angelegenheit.
Und du wirst sehr early auf, dieses...
of an army of smart senior engineers, whatever that will do to the profession of software engineers.
We'll see where it goes, right?
But I think the level of failure just moves up, right?
So your errors or mistakes are happening on a different level.
Not on the code level, but maybe on a different level.
And your plans are not good, and so you need to invest into system design, which means you need to think about system design already as a junior.
If you just start going, you most likely will do something that is not...
Und du lernen es auch wenn es nicht so ist, wenn es nicht so ist, wenn es nicht so ist, wenn es nicht so ist.
from the grilling session that are very technical and along with some other helpful documents like the PRD and such with the help of an LLM that has the head of a specific stakeholder on you then create a document for the stakeholder to actually consume and then be able to answer questions I heard something very similar from two folks actually at a CTO craft event that I attended like a month back or so This seems to be something also, like seems to be a pattern, like using the tools in order to create context for non-technical people.
Yeah.
I also realized, at least in our case, that the PRDs have a different structure system, uses the paragraph symbol and paragraph 6.3, for example, while the technical design documents, it's actually numbered with just 6.3.
So references in...
In dieser sehr technischen ersten Version der Frage, sie waren inintelligente.
Sie mussten mehrere Dokumente openen, aber die erste Reaktion von allen Stakel waren, ich verstehe die Frage nicht.
Und sie mussten mich erst einmal erklären.
Ich dachte, ich mache das jede Woche, das kann ich nicht.
So jetzt die Agent machte die Translation und ich habe, ich denke, noch zwei mehr Reviews open zu sagen, wie es tatsächlich funktioniert.
So ich habe das in.
Let me judge what you're actually doing in order to perfect it a little bit or improve it rather.
And so far the reaction is we're moving not, we used to need like, we used to answer like a question every 10 minutes.
Now we need two minutes per question, which means sometimes it's like, yeah, it's that.
I know exactly what you mean.
And I still have the references so I can go to the document if something is unclear, but it does that for me.
So, essentially, I replaced myself in the session that someone else is explaining it.
I can take it eventually asynchronously offline and I can hand them the document.
I set an ETA deadline.
Please fill it on by then.
I replaced myself.
Yeah, that makes perfect sense.
All right, now let's come to the reality check.
So, we usually ask our guests what are the fucking wow moments that they...
mit den LLM oder AI-agenten?
Ja, du hast du diese?
Well, es gibt wahrscheinlich zwei.
Eine, ich habe schon erwähnt, das RSC, das hat eine RSC, das war eine gute Lehre.
Ich habe eine andere, die mein Frontend Engineer, Madu, mit mir.
Er war einer der slowen Adopters von der Harness.
Er war, ich nicht wirklich die Code, es produziert.
Er hat sich, ja, ich habe diese Issue und es hat sich, ich mache die Nummer ab, es war Multiple Files.
Er hat sich, ich habe drei Lines-Code und diese Issue ist fixiert.
Aber der AI hat sich nicht versteckt und hat sich die Files in order zu fixieren, dass drei Lines-Code drei Lines-Code würde haben.
Er hat sich, was ein Argument und Waterer auf das Melkmal hat, dass er, ja, wir wahrscheinlich nicht wollen es wirklich.
Der Game-Changer für ihn war, er war er mit Fable.
And Fable, for him, that was the breaking point.
He realized there is something.
It's not necessarily AI itself.
There's different models that can do different things.
And I just resigned.
Okay, let it be more expensive because Fable is more expensive.
But if it gets him to actually produce code using the harness and therefore using the standardization that we have built into the harness, that's good for me.
And then, of course, then multiple loops, me trying to fix the pull request that was supposed to make things cheaper and faster.
And I made it so absolutely worse.
I was actually running it and I saw the count going up and other engineers were coming to me.
It's like, hey, you know you did this.
And we observed.
So they waited a little bit because I did the fix and they thought, oh, Jürgen knows what he's doing.
So that was an aha moment and probably also a wake up call for them.
No, I'm probably not necessarily the person that should be the gatekeeper for these pull requests.
It was literally it.
It didn't accomplish anything in two days.
Normally it would have taken 20 minutes.
So I really made it worse.
So looking what you're doing is always a good.
Absolutely.
All right.
Thanks a lot.
And then final question.
Do you have a prediction for us that you want to share with our listeners?
I actually think having observed this a little bit from Toys status to being really, really useful is that whatever is the coolest rage and just three months ago that was loop engineering.
I actually do a regular reality check and I'm listening to AI podcasts and read some blogs and there's always something newer that comes up.
I mean, you take it from loop engineering to graph engineering, which is something that I'm looking forward to in experiment because I think it could actually Make it happen for Endercore, what we actually want to accomplish at a company as an orchestrated marketplace.
We have sub-agents that ultimately own domains and operate within their domains.
But of course, if they operate in isolation, they pre-optimize very quickly to their own success metrics.
But the success metrics is ultimately that we have happy customers.
And how do you abstract that?
How do you measure that?
So you need multiple levels of gates that are interacting with each other, but someone needs to kind of like own the overall direction of everything.
And that we currently do not have.
Even when we create software, it's still very simple in essence because there's one loop running.
And if it is, of course, paralyzable, I have two implementations running at the same time.
They don't know of each other.
They operate completely independently.
Und adding another layer of complexity, I mean, this is ultimately evolution, so I'm back where I started in my career as a biochemist, evolution only allows you to create a more complex system if you actually have an advantage.
So it is on us to actually make that advantage happen, that the outcome of the code is actually measurably better.
Otherwise, the simple system will win.
And that would be something that I'm very interested in.
that I want to actually experiment here at Endercore, make it happen.
I'm currently recruiting folks that are mostly interested in contributing to the AI layer.
We have the backend and the migration of the backend to Java microservices.
That's the Nearshore team.
But even there, I have folks that have an AI background where I want to hire people to actually help me experiment to ultimately create something that is smarter than what we currently have.
And that definitely is agents controlling other agents.
Und wie viele Leer Sie brauchen, das betrifft auf die Capabelle, die Success Matrix und am Ende des Tages auf die Budget, weil es nicht mehr ist.
Definitely.
Vielen Dank, Jürgen.
Es war ein Pfeuer.
Thanks für mich.
Danke.
So, jetzt zu den Recap.
Was hat mit uns gefallen?
Ja, für mich?
The most noteworthy thing was that they are clearly past the review bottleneck, basically.
So their loop-based approach revealed rather the next bottleneck in the planning.
And that's what they are currently working on, right?
And from what it sounded like, the loop-based approach seems to be very impressive, seems to be really solid and producing solid software.
I'm really curious to see where this is going.
Yeah.
Und ich würde auch sagen, dass du Ende musst, du musst du starten fixieren.
So, starten oder schiften, das absolut macht Sinn, weil du auch einfach sagen kannst, schiffst in, schiffst aus.
So, wie kann du das Ergebnis nicht behalten, wenn du nicht in der Beginn hast, in der Speck zu verhalten?
actually a bit obvious, honestly.
We talked about the review topics so often, but that is actually also a good segue to my recap.
For me, I need to think about that longer, but I think what we are seeing in all the interviews is there is a pattern that, as in tech manager or tech leader, you need to have a deep understanding of the software development lifecycle.
And yes, also that sounds obvious.
Honestly, if I look back, a lot of stuff I did during the last years was also on organization level, like people management, all that stuff.
Less about the specific software development lifecycle because that was set.
But in the agentic era, this is now so crucial.
Without that knowledge, you will have a hard time being a tech manager.
This is one of the core requirements.
Yes, I tend to agree.
What sparks now in me is that what has been hard, or at least what seems to be hard, is usually taking the flow of the new life cycle that you created and taking it from One engineer to a team.
He explained that he did this or they did this with their harness.
So they have a harness that provides this lifecycle, provides this infrastructure for their flow, for the loops.
And he has these integration points in order to improve the plans with the other stakeholders where then the questions need to be answered that are coming up during the grill me sessions but yeah the harness really as the the main tool in order to distribute or scale the the agentic software development life cycle that's very interesting that's it for today thank you and hope to see you next time bye bye bye the beyond vibe coding podcast is a project by sebastian heidemar zarten and andre neubauer Partnering with ImpalaSearch.
The content is created by us and our guests.
Join the discussion on LinkedIn or visit our website where we publish all episodes.
For questions and inquiries, feel free to reach out via LinkedIn.
Thank you for your time and see you in the next episode.
