# Sovereign AI Infrastructure and Open-Weight Model Strategy

**Podcast:** HMZE
**Published:** 2026-07-30

## Transcript

Und dann hast du das Geld gekauft.
Ist es eine bestimmte Grund, warum du das Geld gekauft hast?
Das ist eine gute Frage.
Ich gebe meine Best.
Ich gebe meine Best.
I'm Sebastian Heidemeyer zu Erpen, CTO at NorthIO.
And I'm Andrei, CTPO at Trusted Shops.
Great to have you back.
This time we are talking again to someone from ZipGate who are very actively utilizing AI.
Yes, and this is actually the first episode about local AI that we're running.
We have planned more episodes to dive deeper into this topic.
But for now, we hope you enjoy this first one.
Welcome Stefan, we're super happy to have you on the show.
You're the second person from ZipGate actually we're having on the show and maybe we touch upon that in a minute or so.
But before that, please as always introduce yourself to our listeners.
Sure, my name is Stefan.
I have been at ZipGate for 18 years now.
I entered the industry right from university during the dot-com years.
Weil es so exciting war, dass ich aus der Universität aus der Schule angefangen habe.
Und dann bin ich ZipGate ein bisschen später in ZipGate.
Und habe alles gemacht, was alles ein tech-personer kann.
Ich begann als Frontend-Developer.
Wir waren der zweiten Scrum-Team bei ZipGate.
Dann bin ich Product Owner.
Can you elaborate what that means?
Making sense of AI for ZipGate or being the AI lead at ZipGate?
Yeah, I can try.
Well, making sense of AI is something we all try, I guess, and no one would say we are able to do it.
So I try to do it in a way that makes sense for ZipGate, which makes it a little bit easier, but just a little bit easier.
As a company, Sipcat has always moved fast.
We were very early in the agile movement.
We have tried to keep pace with everything.
In our industry, I think we were the first ones to experiment with AI.
And now we are in a good position and we were able to start agents from within.
That is the starting point.
And I think 2018 maybe was our first.
Wir waren gefragt, ob wir uns erstens eine Art von AI integriert.
Wir waren gefragt, ob wir uns eine Art von AI integriert haben.
Und die Antwort nach einem Jahr von Arbeit war, dass es sehr interessant ist, aber die Technologie ist noch nicht da.
Wir können nicht verstehen, was die Leute sagen können und wir können nicht verstehen, besonders in Deutsch.
Wir immer nachdenken, und wir wollten das immer nachhaltig, aber wir waren für einen Punkt, wenn es eine Technologie funktioniert.
Und dann, wie alle wissen, die LLM-Revolution kam.
Even Text-to-Text wurde besser, und wir haben angefangen, zu bauen Produkte um es.
During that time, it became clear that building products around it was not the only thing because the building itself was becoming a thing driven by AI.
So engineering changed profoundly in the last two years.
And my job right now is to make sense of how we integrate AI into our daily work, how we integrate it into our products, and how...
our APIs play into that whole field.
Because we believe that very shortly in the future many people are not going to be able to click around and use a product anymore because they will have an agent that does it for them.
A colleague of mine just a few weeks ago said, well, I don't know how to click anymore.
I need an agent to do it.
So if you don't have an agent for it, I can't use it.
I thought that was a very interesting point.
And that opened the eyes for many, many people at Zipgate.
So we are now actively pursuing becoming an, well, AI first sounds like a buzzword for managers who don't know what they're doing, but we still try to become an AI first company that does engineering with AI and does products that use AI to make the product better and not just for having AI in the list of features.
Das ist ein sehr vieles Ziel.
Das ist das ich, was ich.
Ich versuche, es in die richtige Richtung zu integrieren.
Ich versuche, dass wir nicht weg von der Action passiert.
Und ich versuche, dass wir uns in die richtige Richtung mit AI-usage sind.
Das führt uns zu heute.
Well, actually, it was not a practical problem, but it was something we thought about was what happens if we lose access to the technology.
How come?
How did this question come up?
So we'll get into that in a minute.
But before that, I want to ask maybe one more question on specifically your AI product.
oder wie du AI in deinem Produkten verursachen?
Ich könnte dann, dass in der Telefonie-Tex-Tex, ich glaube, du hast das, das ist wichtig, aber nicht nur das.
Du hast auch eine Grunde für die Informationen, die die Agents wollen, oder wie wir das?
Wie kann man das?
Well, eigentlich, alle in der Industrie ist die gleiche Sache.
Sie starten mit Text-to-Text, dann haben sie LLM und dann haben Text-to-Speech.
Das ist sehr einfach so einfach, aber das ist die Stack wir alle haben.
Interessant ist, wo wir unsere Technologie haben?
Wir sind in der sehr lucky position, dass wir vor dem großen AI-Hype haben, wir haben die neue Plattform in einem Weg zu integrieren.
any technology at any point in the whole telephony framework, starting from the phone numbers to the point where we have interconnects to other networks.
So that made it easy for us to integrate those technologies into what we have.
And that is a huge asset for us because we are now able to experiment wherever we want.
We have access to data because we asked our customers very early, can we use We started to anonymize data from our customers who gave permission very early on.
So now we have a full stack of every layer of the technology where we can attach what we do.
And that makes us able to have a very, very quick development cycle and to see where interesting points are.
I often say, Es ist sehr schwer für jemanden, die nicht die Insights haben, zu haben gute Ideen, die die Ideen der anderen zu überlegen.
Und ich glaube, die vollständige Bildung gibt uns eine unique Art von der Sicht, wo wir die ganze Technologie stacken können, und wir können sehen, wo wir neue Technologie anstellen, dass nicht so viele Leute haben gedacht.
Das ist ziemlich unigke mit der Firma.
Wir haben ein sehr lange Zeit investiert.
Wir hatten eine AI-team, glaube ich, für zehn Jahre, das ist jetzt größer als es ein paar Jahre alt war.
Aber noch immer noch Leute waren paid für die AI-team.
Und das hilft viel.
Und zu haben, die alle technologischen Läden und die Daten ist ein anderer großer Asset.
Und vielleicht können Sie ein konkretes Beispiel aus dem, wie das in einem konkreten ...
Let me think about which one I can tell you about.
They did that very early on for a pretty simple feature, actually, which is like we do call analysis.
So for this call analysis, we used a huge LLM like GPT for a long time.
And we had all this data, all the data coming in and coming out of the LLM.
So we thought, why do we need that huge LLM with a huge prompt that tells it what to do?
Wir können einfach nur die LLM ausführen, um die Daten wir brauchen.
Wir konnten die Required Compute für die gleichen, wenn nicht even bessere, und das war, weil wir die Daten in die LLMs hatten.
Und als ein Side-Effect, wir können das LLM auch hosten.
Das macht die Daten immer mehr sicherer als es vorher war.
Das ist eine Sache, die wir able zu tun haben.
Ein weiteres Thema ist, dass wir viel Daten von realen Callen zwischen Menschen haben.
LLMs sind oft auf Wikipedia-Buch und Sachen.
LLMs sind nicht so, wie die Leute sprechen auf dem Telefon.
In der besten Fall, dass sie wie die Leute sprechen auf dem Telefon in den Film und Buch sprechen können.
Ja, sehr anders auf dem Telefon.
Und ich denke, wir können wie ein Computer in LLM sprechen zu Menschen auf der Telefonklinie in einer Art, die sie fühlen, wie sie sich mehr natural und wie ein Mensch.
Interessant.
Ich glaube, dass das nicht nur über den Konten ist, sondern auch über den Text-to-Speech-tonal, wie die Tonalität, die pausen und alles.
Wir arbeiten mit Partnern für das, weil es ein viel größer Daten-Projektions-Hurden für Voice Audio-Daten ist für Text-Daten.
Das ist etwas, wir denken über, wie wir es tun, aber wir wollen nichts falsch machen und nicht mehr machen.
Wir müssen nicht alles falsch machen, so einfach Dinge wie Anonymizing das Daten.
Es ist eigentlich, heute mit der Technologie, wir können eine Modell auf die Person trainieren, dann können wir die Person verwendet werden, mit der Person mit der Person mit der Person mit der Person mit der Person mit der Person mit der Person trainieren, dann können wir das Daten trainieren, aber das gibt eine große Margin für Errungen.
So diese Dinge sind kompliziert.
Für sicher.
Eine Frage.
Sind Sie dann auch wirklich mit der Modell trainieren?
Oder so, zu welchen Grad?
you're using AI?
Is it training?
Is it fine-tuning?
Is it just inference?
Well, we are experimenting with training, but right now I don't believe that training will make what we do better.
I think fine-tuning is the way for us to go because training can be very expensive and the base models are actually pretty good, especially the small ones.
Und du kannst viel mit einfachen mit dem.
Ja, awesome.
Thanks a lot.
So, das war ein bisschen länger als erwartet, aber es war interessant.
Das ist warum wir diese Frage stellen.
Und bevor wir zu der Hauptsache, die LLM-Soveranity, eine kurze Frage, wie du jetzt gerade arbeiten, was ist dein Tech Stack?
Wie soll du das AI-utilizeieren?
Ich denke, ich benutze die meisten LLMs jeden Tag, aber ich benutze Cloud die meisten.
Und als ich aus dem Entwicklungskontext habe, ich benutze Cloud Code für alles.
Also für schreiben, documents, preparing meetings, ich benutze Cloud Code.
Wir haben einen internen Agenten, Aronux, die ist able to...
access to a certain extent internal systems to get some data.
I use that a lot to get statistics, things like that from our systems.
But my main workhorse is just the Cloud Code command line tool.
And I have some projects.
Most of them are documented in Markdown.
I use Obsidian to have a better view of what's going on.
And Cloud Code is doing a lot of the heavy lifting, actually.
I try to get transcripts of all the meetings I'm in, and I try to get my stack to be clever about putting these into perspective and putting them into the right folders.
But inside the company, we work like...
We throw over a lot of clawed artifacts to other people and work together on those things.
I think the workflow went from a lot of text files, from a lot of Google Docs to a lot of HTML.
Most things have these, most things, most artifacts we share in the company look like clawed code built artifacts now.
I think you know what I mean.
Ja, ich benutze viele, weil ich nicht in Cloud habe, weil ich alles für alles habe.
Aber mein größter Punkt ist, dass ich eine große Rolle mit dem Strategie habe.
Ich weiß Markus Andrezak, und er hat uns gesagt, dass Strategie jetzt Infrastruktur ist.
Und ich denke, das ist eine der Hauptpunkte, wir sind nachher, weil nur Kontext macht die Resultate gut.
Und meistens die Menschen sind da, die Kontext zu geben, aber wir sind glücklich, wenn die LLM kann, wie viel Kontext aus dem Structured-Data bekommen.
Das ist das, was wir versuchen.
Und dann haben wir Automated agents are running.
They're all running on our Aranax system.
It's that internal AI that schedules crons.
Usually those things are not AI.
In the beginning, everyone was like, yeah, we need to build AI systems.
And now it's more like build a Python script that fetches the data because especially with these daily reports, it's always the same.
There's nothing an AI could bring that.
We can't do it with Python or stuff like that.
Usually these reports are not very intelligent, but built by AI.
I think that's one thing where a lot of this is going.
I mean, you build a dashboard with AI, but you don't leave the interpretation to AI in many, many ways.
Although we sometimes try to have AI find new facts, new information in the data.
mit neuen Insights.
Aber mit Existing Produkten, die zehn Jahre lang sind, finden neue Insights für ein AI ist nicht so easy.
Ja, ich kann es nicht.
Wir haben eine Phone-Kompanie, so wir haben eine Menge Fraud.
Und Frauders werden immer cleverer, so sie wissen, wie sie können, wie sie ihre Sachen machen.
Sie bauen Webseiten, die sich als ein Unternehmen ansehen, Just for us, when we check if they really exist and stuff like that.
And that works very good with AI.
So AI is very good in recognizing frauders and stuff like that.
We like that a lot.
I think that's many hours per day that we can pass on to an LLM.
You mentioned Aeronux.
Is that a custom thing you're using to give a bit of harness to the entire organization?
Ich würde sagen, es ist die Open Claw von ZipGate.
Es gibt nicht viele Möglichkeiten, weil wir uns zu retten und zu retten, weil wir die Daten haben, aber es ist ein sehr guter Tool.
Es hat eine Art von unseren Insights zu haben und kann sie mit dem, um mich zu machen, wenn ich eine neue Frage habe, das nicht mehr kommt.
Wenn ich eine alte Frage habe, dann einfach nur die Daten in die Artifacts finden.
Cool, danke.
Jetzt ist es Zeit für den Meet, für den Hauptsame Thema.
Du hast bereits hintet, dass du schon ein Stück selbst selbst schon seitens, also ein kleiner Models.
Aber ich denke, da war ein LinkedIn-Dieter von dir, wo du ein Experiment von dir beschrieben hast, dass du ran hast.
Maybe that's a good starting point.
Okay, that's very exciting.
We've been thinking about, is there a reason to buy hardware for a while?
We weren't really sure.
We had companies make us offers, but we were not sure if we were going to do it.
But then the Claude situation happened where where the newest Cloud version was kept from the customers, not just in Europe, but officially just from people from outside the US.
So we got into thinking and pretty quickly we were sure that we needed something to mitigate for a situation where maybe even older models will be locked for European companies.
So what we did is we rented Pretty big machine, Blackwell Architecture B300 machines from Finland, and installed the, at that time, best, biggest coding model available as OpenWeights, which was GLM 5.2.
And we just reached out to all of our developers and told them for a week, use that.
Stop using Cloud Code, I think.
About 90% of our developers are using CloudCode.
Some are using Codex.
Some are using local models for smaller stuff.
But most are using CloudCode.
And we told them, try this.
And they didn't have a harness.
Some even tried to use it inside the CloudCode command line app.
Some used PyAgents and some used OpenCode.
And all of a sudden, we had a lot of traffic on that one machine that was running the GLM model.
At one point, I saw about 80 sessions at once.
So that's actually 80 requests to the LLM running at once.
And the machine didn't even blink an eye.
The verdict of most developers was, this is so much faster than Cloud Code, which was impressive.
Although the machine was kind of expensive too.
7,000 euros a week is something.
Das ist nicht so gut, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es so, wenn man es auf einer der Unternehmen, die sie bei Token-Use verkauft.
Aber jetzt sind wir auf den Hardware in-House und installieren diese Models, die wahrscheinlich nicht GLM sind, aber Kimi, die ist viel größer.
Wir sind ein bisschen Angst, dass es nicht auf diese Brand-Newen Machines fit ist, aber ich habe es gehört, dass sie funktionieren.
Und wir werden sie in unser eigenes Datacenter setzen.
And then our developers will probably all work on that.
We will not tell them never to use anything else because we believe that experimenting, trying new stuff is very, very important.
But for us at the moment, especially with coding, it's not just about data protection.
It's not even a very important thing because, I mean, we have our code on GitHub.
So that's already there.
So the most important thing for us is sovereignty, access to our own hardware, access to models we want to use and not being cut off of the action when we really need it.
Adjusted the mass, you say it's 7,000 Euro on a weekly basis, like 50 weeks, like 350,000 Euro.
Es ist viel, aber wenn man die Masse auf 100 Cloud Code abonnieren kann, dann über die Überusage werden, dann wird es in die Richtung.
Es ist nicht komplett off.
Du könnt vielleicht auch noch mehr Geld geben.
Du hast gesagt, es war nicht fully genutzt, das Maschine.
Ich denke, du könnt es auf ein paar Mal aufpassen.
um, like seven figures, um, we would pay for Entropic.
Yeah, yeah.
So, and if I saw, then like even better, right?
So the use case is not on, uh, saving money or like spending more actually.
Um, it's, uh, even that case is positive.
Yeah, yeah, absolutely.
Everything, everything is like, um, there is, um, Almost nothing that gets worse, actually, which is not something we expected.
I mean, I was like, my workflow when I write code, when I write code, I mean, I'm not a developer anymore, but I still know what's happening there.
I usually just take screenshots, give them to Cloud Code and tell it, yeah, here, that looks bad.
Do something about it.
But we figured that most of our developers don't do that usually.
They are more text-based.
Most of them just speak into their microphones and tell Claude what to do.
So GLM doesn't have vision.
So that was my biggest fear, that they would not be able to follow with their workflows anymore.
But most developers don't have that workflow, actually, I found out.
So most just talk into the microphones and have the LLM do what they do.
So there's really not much that got worse.
I think most developers were pretty sure that the quality is Almost on par, more like Opus 4.6 maybe, but still pretty good.
And everything we heard about Kimi, and I already tried Kimi on the Moonshot AI servers, it gets even better.
So we're really excited about it.
Yeah, if you have a good harness in place for software engineering, which allows you to replace models, I think that should be then also a win-win-win.
I would question quality.
But as you just mentioned, even there it was on par.
Yeah, I mean, it's hard to measure software quality on a very quantitative level.
But as far as we could see, quality was very good.
And in some cases, developers said the quality was even better than what they would have expected from Claude.
That's interesting, because I would have assumed that The engineers would be very skeptical towards the model, which I'm not sure if this is true in your case, but if they are skeptical, then hearing generally quality is on par or even better.
I mean, this would be the biggest blocker from my side.
Like if you don't run proper evals where you check if in your cases the model works better or worse or whatever, right?
Das ist sehr wichtig.
Kind of easy, weil sie mit ihnen zu sprechen und sie können helfen.
Aber es gab immer eine Anwesenheit für neue Technologie.
Aber in diesem Fall, da haben viele promises gemacht über Models, die so gut sind als etwas.
Ich sage immer, wenn du zu einem Trade Fair und du mit jemanden gehen, dann wird jeder sagen, wir haben unsere eigenen Models.
Ich bin sicher, dass ich jemanden noch nie gefunden habe, der wirklich hat ihre eigenen Models.
Und da sind so viele Lies, dass ich das nicht so gut expected, aber es hat sich nicht funktioniert.
Aber es hat sich wirklich toll.
Und einige Leute kommen mir und sagen, wenn wir das wieder benutzen können, weil ich es so gerne mag.
Ist es eine bestimmte Grund, warum du das Experiment stopped?
Du hast gesagt, dass du viel mehr spendstatt hast, so es ist comparativ cheap.
Und dann hast du gesagt, dass du das Geld hast.
Ist es eine bestimmte Grund, warum du das Geld hast?
Das ist eine gute Frage.
Ich gebe meine Best.
Wir haben die Best-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-Den-D Wir haben nicht die besten Prices wir für die Maschinen bekommen, weil wir andere Sachen aus dem wir nicht machen.
Das ist das in der Fall.
Und ich glaube, wir haben Glück, weil sie uns gezwungen, die Maschinen für uns gezwungen, weil sie uns eine lange-term Kontrakt haben und gesagt haben, wir werden sie für uns gezwungen.
Und ich glaube, sie hat das für lange Zeit gemacht.
Und eigentlich, wir wollen die Harness richtig.
Having everyone on Cloud Code is fine, because everyone's going to use it somehow in a similar way.
And we want to make sure that we have a harness for the open models that works for the whole company so that we get consistent results.
And I don't want to handle situations where maybe a team or several teams have worse results than others.
And we just want to make sure that we have a consistent...
Harness, a consistent technology in the background.
So, yeah, we don't have to handle edge cases that wouldn't appear otherwise.
Yeah, that's what you mentioned, right?
That some folks didn't actually have harnesses before because they were using Cloud Code, so they needed to switch to something new.
But honestly, so I'm a big fan of Pi.
I've been saying this on this podcast.
Ich habe verschiedene Modellen versucht, wir haben ein ähnliches Experiment auf einer kleinen Maschine.
Wir haben einen DGX mit 8 A100-GPUs, das nur 40 GB per GPU oder so.
Aber wir haben die QN 3.6, 35 Billion.
Das war ein paar Maler-Parameter-Modell, mit einem Mixer-Experts, immer 3 Billionen aktiv.
Und wir haben einen anderen Experiment.
So LMperf, 16 users, und der Durchput war einfach nur auf.
So für den Test-Duration, wir nicht saturate.
Wir haben nur vier GPUs, wir haben nicht saturate.
Und honestly, auch die Qualität der kleinen Modell war sehr, So compared to probably anything a year ago, this was already quite a lot better, including reasoning, right?
And that's given the fact that I'm using a harness that has a lot of customization, that has skills, that has the code base, right?
There's a lot of grounding for the LLMN.
And if it does then have thinking capabilities, then that...
ist ein ziemliches Problem.
Right now, ich arbeite mit Ornith.
Ich denke, das ist das, was ich test auf meinem Local Machine, die Optimized Coding Model, based auf die Alibaba QIN 3.6 Model.
Und ich habe zu sagen, dass das ist, dass in der Falle der Frontier Labs nicht available ist, Das ist eine gute Lösung, das nicht mehr zu einem Bauen, sondern eher zu 90% oder 95%.
Ja, ich denke, eine der Dinge, die wir in der Harness zu bauen haben, ist, dass wir besser entscheiden, was Models zu nutzen, was für welche Use Case.
Das ist ein hoher Thema.
in the scene.
And I guess we are going to have to do that, too, because sometimes even a fine-tuned model for a specific case might even be better to solve a problem than a large model is, even for coding.
And so, yeah, we'll have to look into that.
That is also one of the things we're definitely going to look at.
I have on my desk at work, I have one of these Studio Max M8 Ultra, biggest box they ever made things.
And that's a lot of fun to play around with the models on that machine, actually.
And it does even run GLM because it has 512 gigabytes of RAM.
Okay, not that bad.
That's very slow.
That would be my follow-up question.
So the hardware you're just buying, it is for coding?
First question and the second is, have you looked into local hardware?
So local hardware, I mean like workspace hardware.
You just mentioned MacStudio.
So I'm wondering why this could be a pass too.
I believe it can be a pass.
But right now, I think for what we're doing with understanding large code bases Ich denke, es ist gut für uns zu haben, dass wir zu einem großen Modell haben.
Ich bin nicht sicher, ob die Modells sind, oder größer.
Wenn wir die Beziehung für die Machines entschieden haben, ein unserer Founder hat gesagt, Ja, ich weiß nicht, vielleicht ist es zu groß.
Es ist alles zu klein.
Wir brauchen die große Maschinen.
Und nach Kimi kam aus, er kam zu der Office und kam mit der CTO und sagte ihm, wir brauchen largerer Boxen, weil die Models sind größer.
Wir wissen nicht, ob die Models sind größer.
Ich meine, da ist vieles research.
Du kannst 38B Models installieren auf iPhones mit 1-Bit Quantenzeichen.
So, maybe it's all getting smaller, maybe it's all getting bigger.
But one thing that at least the Studio Mac has is it's always kind of slow compared to other enterprise-grade hardware.
I mean, even compared to consumer graphics cards, it's kind of slow, although those are so expensive you can buy them.
So, I don't know.
We didn't play with these 5,000 Euro machines, AI machines, that everyone's building right now.
We didn't try those, but I think for us right now, a centralized solution is better.
We didn't specifically buy the hardware for coding, actually, or decide to buy it for coding.
We thought...
A lot of our other models can run on that hardware too.
For our customer-facing products, we already run our own text-to-speech and speech-to-text models.
Okay, I'm also talking about our own models.
They're not our own.
Other people's models on our own hardware.
And also have an LLM that we can use for that.
So it's a smaller one though.
There is an aspect of using it for our customer-facing products, but we are also thinking about buying just a little bit cheaper hardware for that.
I'm sorry.
Sorry.
Buying cheaper hardware for that.
So, yeah, I don't know.
I mean, the thing that makes calculating the cost efficiency the easiest was coding.
Making the decision based on money was the easiest to just say coding is much cheaper.
Absolutely.
And I also think that there are differences, right?
So on the one hand, if you want to replace the infrastructure for 80 or 100 engineers, then that's a different thing than having something for one person to do some use cases, right?
To utilize some use cases.
And this is where I would also differentiate, specifically also given the fact that usually for coding, you have huge contexts that you always need to transfer, right?
So you always have huge context windows that you need to have.
And then the model itself plus the context size, all of that needs to be readily available.
And also they're quite like these complex coding agent.
There's constant turns between the agent and the GPU server.
So they are quite compute-heavy and extensive.
That's a totally different ballgame to a chat assistant or a simple rag or something like that.
Absolutely.
But that actually also translates into if you're a larger engineering organization or mid to large engineering organization.
You should always go for central infrastructure.
True?
For now.
Only for now.
We are not in the section where we do any kind of predictions.
Yeah, exactly.
I think, well, talking about predictions, I believe that the American Das ist nicht mehr so, dass die EU nicht einfach nur die Öffentlichkeit von China auswählen.
Ich hoffe, dass die EU nicht einfach nur blinde und dass wir noch immer noch die Öffentlichkeit von China auswählen können.
Wenn das der Fall ist, und sie developen in der Zeit, dass sie jetzt gerade jetzt sind, Ich würde sagen, dass es ein paar Monate später mehr kost-efficierter ist, dass es Ihr eigenes Models ist.
Und Sie könnten auch sagen, dass Sie nicht nur die Infrastruktur auswählen, sondern auch die EU-basierten Models auswählen?
Also haben Sie also looked into, was Mistral ist das?
Will das sein?
Was ist Ihr strategy in der Trugout?
In der Produkte, heute, kann man zwischen einem American-Modell von OpenAI und einem Mistral-Modell.
Although ich weiß nicht, wie die Sovereinheit von Mistral-Modells hält, es ist immer etwas, was unsere Kunden wollen.
Ich denke, ein ein-figure Prozent der Unternehmen ist, die nicht über die Software-as-a-Service-Service zu gehen.
Ich denke, das ist ein großer Punkt.
Und es ist auch ein großer Punkt.
Ich meine, ZipGate, historisch, wir haben nicht große Unternehmen, sondern wir haben nur ein kleines, kleines, medium-sized Unternehmen.
Also, sie sind über das, viele von ihnen.
Wir haben viele Leute, viele Leute aus den Medizinischen Fällen.
Wir müssen sicher sein, dass unsere Daten ist in einem Weg.
Das ist etwas, was wir definitiv haben, und als Unternehmen, wir wollen sicher, dass niemand in diesem Traum fällt.
To say it carefully, many people from professions where they would have a very big incentive to have good data protection are not very careful with their data.
So for us, it's very important to give them defaults that protect them.
And so I expect European hosted or even hosted on our servers models to be the default by the end of the year for everything.
That's a word.
Ja, aber ich kann es auch mit dem, ich fühle mich auch ein starkes push in dieser Richtung oder Momentum in dieser Richtung und natürlich, der Fable 5 incident war ein großer Böster und gegeben, am ich an dem gleichen Zeit, diese Chinese models being on par oder partially even better looking at Gimmi K3, which ist jetzt, ich denke, die LM-Archstellerin, wenn es um Coding geht.
Ja, so warum nicht?
Ich denke, es ist total total exciting.
Ich meine, um, um, die Models zu Hause und die Harness zu kontrollieren, gibt es so viele, um, wirklich exciting ways zu arbeiten mit LLMs und zu offeren, um, inside der Firma zu unseren Kunden zu erzeugen, um, dass ich wirklich über die Zukunft freue mich.
So, um, wir können wir auf die Pace, um, So many Software-as-a-Service products, which I heard are getting worse at the moment because people are just vibe-coding stuff in without proper harness and without thinking about what's going to happen.
I think Software-as-a-Service is getting so much better in the future and so much faster development cycles.
And that's really exciting to me as someone who came from the .com era.
That's a very positive outlook, yeah.
One final word also on the model, like bigger versus smaller discussion, right?
Something I haven't managed yet is try out this, I think, 1B mini CPM model that has been fine-tuned using Claude Fable 5 traces, which has been at least advertised as your mini Fable 5 on your own machine.
Try to test models like this.
Because if this 657 MB model would actually be on a similar quality level, then this would really flip the big versus small model discussion quite a bit.
Absolutely.
But I always like what's going in my mind when I think about these is more like this.
We have stuff at home.
We have hamburgers at home and you get this.
German bread with just a piece of meat in between.
So, right now, I was always a little bit disappointed with smaller distilled models.
But I think the way, way more interesting thing is specialized models for certain things.
Like a specialized model for unit testing, a specialized model for, I don't know, user experience design.
Those things might be Das kann ich sehr cool für den nächsten Mal sein.
Und Sie können Distill.
Ich kann Fable-Skills in Unit Testing in ein Modell und ich glaube, ein kleiner Modell wird viel besser als ein kleiner Modell, wo man all die Fable 5 in den Modell tryst zu kramen.
Für mich ist es ein bisschen die Frage, was passiert, wenn man ein Modell distillt, wenn man ein Modell macht, man kann es dann auch ein Teil der Modell und man kann es dann auch ein anderes Modell.
Allerdings zu was du, was du gesagt hast, du hast den User Experience Testen oder den Unit Testen, was du einfach nur die Werteilung für das.
Die Frage ist wirklich, wie dense ist die Werteilung in diese huge Models in der Realität?
So, wenn du es, ist es ein großer Werteilung der Werteilung, das ist eigentlich nicht 90%, 95%, maybe 99% even, because it's essentially all the knowledge of humanity, if you wish, that you have in these models.
And then wouldn't it actually be possible to distill a 1%, 5%, 10% size model that contains most relevant knowledge for 99% of all use cases that you can think of.
This is the other discussion, but anyhow, this is a very theoretical discussion.
And then you would have to have a very concrete idea of the different knowledge areas to pretty much generate the learning path for the distilled model so that you are building up on the very basic knowledge that is there in the beginning.
after a basic training, right, in order to then have it learn specialized knowledge areas, if you wish.
But I'm pretty sure that this is an area where there will also definitely happen a lot over the course of the next 6 to 12 months.
And we'll see many smaller models perform at great quality levels.
Like for me, the iOpener was really the QEN 3.635B model, which In meinen testen, ich habe nicht gesehen, wie groß die GBT 5.5 ist.
Aber ich habe nicht gesehen, dass ich eine große Unterschiede habe.
Wenn das Modell war, was nur auf Coding und nicht auf Word-Knowledge-Tür, dann würde ich vielleicht sogar besser werden.
Also, für ein guter Developer, du musst immer wissen, dass du ein paar Sachen rund um.
Das ist nicht ein development.
Ich wollte das einfach mal sagen.
Ja.
In der Zeit, wenn wir die Visiten, besonders Press in der Office hatten, wir haben noch viel Art in der Office.
Und der CEO immer gesagt hat, wie jemand, die nicht gut ausfällig sind, ist es zu machen, gute Art zu machen?
Ich meine, ich verstehe LLMs nicht gut, wenn das eine Influenz hat oder nicht.
Aber ich glaube, es sollte.
Ja, danke.
Ich denke, wir sind in der Ende dieser Episode.
Wir haben sehr viel über die Sören gesprochen.
LLMs and really interesting insights from your side.
We have two smaller segments towards the end of the episode always and one is the reality check.
So what are wow moments or even maybe what the fuck moments where you still believe so wow we are doing all of this stuff but this is not working.
Why is that?
So do you have these kinds of moments?
Yeah I'm going to go back like back to when LLMs were very very new and Wir waren einfach nur auf was sie waren.
Und an diesem Punkt war ich mir über einen Experimenten zu tun.
Und ich war mir, was wenn die LLM könnte jemanden etwas über das und sie würde glauben, dass die LLM das nicht wissen, ohne es zu wissen.
On the carnival, the people who tell you your future, how do they do it?
And there are some techniques to do it.
So I put this technique in the prompt.
Back then, models had no vision.
There was no vision model at all.
So I took one of those models that tags photos with what's in them.
That was a little bit more detailed.
It could tell you there's a person with a beard and glasses.
So I used that.
And then I took a picture of people.
Ich pute das in die LLM mit dem Technik, die Leute benutzen, um die Fortunatel zu erzählen.
Ah, Fortunatelers.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunatelers, ja.
Fortunat Es war, dass ihr Zukunft blieb oder so etwas.
Es macht eine Prediction.
Und es war eine Person in der Firma, die gesagt, dass ihr Zukunft nicht gut ist.
Und ich war, ja, das ist einfach random und es geht nicht so.
Oft sind wir die erste Person, wenn das passiert.
Und die Person war in der ersten Seite, der zweite Person war, und es sagte es wieder.
Wow, ich hoffe, es war wirklich ein Fortun.
Ich denke, es war ein bisschen concerned.
Das war das erste große wow-moment, wo ich gesagt habe, das Technologie ist crazy und Dinge werden aus dem es kommen.
Ich denke, es ist auch ein Metapher für wie wir LLMs sehen und was wir denken, was wir können, was wir vielleicht nicht machen.
Ich bin sehr sicher, dass es nicht mehr Dinge über Menschen wissen.
Aber wirklich, das ist ein sehr guter Punkt.
Ja, das war die größte Wow Moment.
Ich habe eigentlich einen Moment, wo ich sagen habe, ich habe eine sehr gute Graspung auf was sie können und was sie nicht.
Ich habe nicht gewusst, wenn etwas nicht funktioniert.
Es ist mehr so, ja, natürlich.
So you try it out thinking maybe it could work, but assuming it won't work, right?
And then it doesn't.
Yeah, and it doesn't work and that's okay.
I mean, one thing when you are in a real-time voice that is sometimes a little bit of a lackluster is when models are not able to be fast enough to do what you want to do.
I mean, voice has to be quick and we're working on technologies to make it even quicker, but then You have 100 trials and everything works, and then you present it to someone, and then there's a network problem or anything, and it takes five seconds to answer.
And then the founder of the company comes to you and tells you, well, so when are you going to build the real product that we're going to ship to customers?
And you're like, well, that is exactly the product we're shipping to customers.
Maybe that.
Maybe things that are unexpected, that don't depend on the LLM.
sind diejenigen, die ich manchmal denke, ich bin disappointed.
Ich war viel mehr.
Vielen Dank.
Und dann, finaler Punkt.
Do you have a prediction für unsere listeners?
Ich könnte, ich kann ein Bild und den LLM machen.
Today I talked to our CEO and we were both like, I don't think we know what we don't.
I think we don't know how we are going to work next year at all.
And we have no idea.
And although we are trying to find out how we're going to work next year, we still have no idea.
And so I believe that we don't have an idea, but I believe that it's going to be.
Ich glaube, dass wir uns mehr mit AI arbeiten können, was wir heute machen.
Wir werden uns Menschen finden.
Was sind die Dinge wir können besser als AI machen?
Und wie können wir sie besser machen?
Und ich glaube, dass wir uns besser werden können, weil wir das wissen, was uns ausgibt.
Ich glaube, dass wir uns besser als Dinge machen können.
ourselves better and knowing how to use AI to amplify those things that we can do, we are going to have very cool products and we are going to have, in some cases, things that help us make our lives easier and also our lives so much harder.
I mean, I'm always coming back to the dotcom bubble back then.
We were all enthusiastic.
There was nothing negative about all of this.
Dann kam social media.
Und so, ja, wir wissen, was wird passieren.
Aber ich glaube, wir haben eine Chance, um zu finden, was uns zu helfen, um zu verbessern.
Und ich wirklich glaube das.
Und ich hoffe, es wird das so, nicht das andere.
Und ich glaube, dass wir nicht das so haben, dass wir AGI haben in einem Jahr haben.
Das ist ein einfaches Beispiel, aber ja, danke für das.
Ja, awesome.
Danke für die Erinnerung, und ich habe wirklich gefühlt den Sprechstunden.
Danke für mich.
Gut, jetzt zu den Recap.
Was hat mit uns, Sebastian?
Ja, für mich?
The impressive thing was, number one, I looked it up actually, that you actually need eight B300s to run GLM 5.2, which I just didn't know.
But then what was even more impressive, that with this cluster, they were able to serve 80 engineers, or probably more, faster than Cloud Code.
Really impressive.
I wouldn't have expected this.
Maybe what could be a reason for that, that it's the Claude Harness, right?
So Claude Quote Harness, which is maybe holding back here a bit, but not sure about that.
Not sure, but it's fascinating.
As you said, I think they probably have Also plan that infrastructure, maybe also for more people, even though you're right, that is actually the entry barrier, right?
This is crazy.
Actually, my first learning takeaway is in that same field, practice always beats theory.
So GLM 5.2 is on par with US models for most use cases, which is Ich denke, das ist ein tolles Geist, weil wir alle immer in der höchsten Boxen immer wiederholfen.
Tropic ist, dass Fable alles in den Fable ausfällt.
Und ich denke, für die meisten von den Fällen, kann man auch einen höheren Boxen nehmen, aber man kann auch andere Models in dieser China-Models.
Und mein zweitens Take-Away ist, es geht eigentlich um die Hardware-Töpe.
So, Central Hardware versus Workspace Hardware.
A lot of people are talking about the next Max Studio and M5 or M5 Pro Ultra, I don't know.
So the idea of running Models locally, so not on your own hardware, but really on your own desk, so to say.
I'm not sure after this talk today.
I think if you're at a certain scale, at least budget-wise, you could also afford central hardware and then go for...
like really frontier models rather than having to deal with well-listed, very compressed models just to make it work on limited hardware.
Yes, that's really the tricky thing there, right?
I think a DGX Spark with 128 gigabytes of RAM is great for running it with Gwenn 3.6.
mit 35 Millionen parameters oder sogar mit größeren Modeln, das sollte wirklich funktionieren.
Also, es ist auch okay, aber wenn Sie diese große Modeln wollen, dann Sie brauchen etwas größeres.
Und ja, was ich nicht erwähnt, ist, dass du die 8B300 für die GLM 5.2 in vollen Precision zu machen.
So, wenn Sie die Quanten, dann Sie nicht brauchen diese große Größe von Maschinen.
Das ist also, dass er das auch auf seinem Desktop-Destocken kann.
But the general question that also pretty much or the interesting topic is really like, yeah, are models getting actually bigger or smaller?
So do you always need the biggest models?
I mean, Kimi K3 now is a huge model, also way bigger than GLM 5.2, I think.
And it beats GLM 5.2 and it beats even...
Fable 5 in Benchmarks.
The question is really compared to some other smaller models, where's the sweet spot?
Where's the trade-off?
And that is also super interesting and makes it so exciting.
It's another dimension of change that just...
Und wir wissen nicht.
Ich bin sicher, dass es sich an die Modell Distillation gibt, die Modell arbeiten, auf Laptop, die sind relativ schnell und auch sehr schnell und auch sehr viel für viele Werteilungen.
Und die Frage ist, was die Werteilung ist?
at hand, right?
There are a lot of coding workflows actually that are established that you can use for different harnesses that make use of different model types already.
So you usually utilize a very big model of the frontier model for planning and then you utilize smaller models for implementation or for like code scouting or code reviews or something like that, right?
And this seems to be working quite well and that's a hint that, yeah, differentiation is probably the key here.
And one thing, sorry, I'm now really getting into a talkative mode.
I just yesterday saw a presentation of LM Studio Bionic, I think is the name.
That's a great product for conscious people who want to run their own LLMs, but not for every case.
So it is a tool that you can run on your computer.
It's pretty much like a cloud code or a codex.
So you can upload documents, you can use Text-to-speech with all locally on your own model served by LM Studio and can run use it similarly can build agents that do stuff.
But you can also say for this specific use case, I want to use a cloud based model, obviously then served by the LM Studio Cloud because, okay, they need to make money somehow, right?
But it's a great, like fully fledged solution for people who are conscious.
So maybe that's a hint for everyone who's interested in local AI.
LM Studio Bionic seems to be a pretty interesting tool.
I haven't tested it yet, but I definitely want to test it.
It's maybe a good activity for the week.
And if weather is bad, maybe one thing I just realized, and we also haven't covered that, is there's a huge advantage if you run your own models.
Doesn't matter whether it's central or like workspace hardware.
No question on the data you're putting in, right?
I think that's a huge advantage.
So definitely something where it also makes sense to look into LLM Studio.
Yes.
And with that, I think that's it for today, right?
Thank you and bye-bye.
The Beyond Vibe Coding Podcast is a project by Sebastian Heidemeyer, Zerpen, and André Neubauer, partnering with Impala Search.
The content is created by us and our guests.
Join the discussion on LinkedIn or visit our website where we publish all episodes.
For questions and inquiries, feel free to reach out via LinkedIn.
Thank you for your time and see you in the next episode.
