# Open-Weight AI Models: Shifting from OpEx to CapEx

**Podcast:** Thoughtworks Technology Podcast
**Published:** 2026-09-03

## Transcript

Hello, everybody, and welcome to another edition of the ThoughtWorks Technology Podcast.
My name is Ken McGrage.
I am one of your regular guests.
I'm very happy to have two folks from ThoughtWorks specialize in some of our AI and platform things.
I'll let them introduce themselves.
So first off, Kara Sanders, you want to go ahead?
Of course.
Thank you.
Hello, everyone.
I'm Kara Sanders.
I'm a principal product engineer focusing on AI and agent platforms at ThoughtWorks.
And Andre Almar.
Hello, Ken.
Thanks for having me here.
It's a pleasure.
So my name is Andre.
Right now, I'm working as a lead consultant inside the data and AI service line.
Great.
So we're actually going to want to talk about open weight models today.
It's been something that's been a very hot topic for a variety of reasons that we'll get into, not the least of which are financial as well as, you know, just...
control of your own data and all of those things.
So I think first off, just so we're on the same page, because this is like many terms, very overloaded.
So at least for the purposes of this conversation, and Andre, I'll give you a first shot.
How do you define an open weight model?
What's it even mean?
Sure.
I like to explain this using an analogy.
For example, let's imagine that you sit down at a five-star restaurant.
So you order...
the dish, eat it, enjoy it, but you don't know the exact recipe, right?
So you cannot step into the kitchen and you cannot tweak the dish, right?
So these are what we call the closed-wait model.
So things like...
chat GPT, GPT 4.0 from OpenAI, Cloud 3.5, Sonnet, Opus Faber from Anthropic, Gemini from Google, etc.
So you can only interact with these models via a website or an API.
There is also an open source code.
PyTorch, Megatron by Nvidia, VLLM.
So these are framework software tools and code-based frameworks used to train or run models.
And then specifically talking about open weight models, this would be the pre-cooked dish plus the full recipe, right?
So the original chef of the restaurant that you just ate, spent millions of dollars and weeks cooking a massive meal.
So they package up, they finish the dish and hand it to you.
But they hand it to you along with the exact step-by-step recipe and ingredients list, right?
So the weights.
So you can eat it right away.
You can host it in your own house.
You can toss in extra spices, change the flavor, whatever you want.
So the real world examples on that bring it to the world of AI and machine learning.
It will be models like DeepSeq, GLM, that Keras knows a lot about it, Lama from ETA, Quinn from Alibaba, and so on and so forth.
So is it all financial?
Is it access to the data?
Is it tuning?
Is it all of the above?
You can, for example, using the open weight models, you can download the file, the model containing the billions of weight numbers, right?
Because they are all trained.
And so you have models with different parameters, like 30 billion parameters, 7 billion, 20 billion, you name it.
So you can run it.
You can use those open weight models to run it in your own server.
like Mac Minis are hot today because of that.
You can run in your own laptops without needing permission from the original creator.
So you cannot take a closed-weight model like Fable or Sonnet from Anthropic or GPT-40 5.5 Sol from OpenAI and run it in your own machines, right?
But with open-weight models, you can do that.
No problem.
So...
I think I know the answer, but just for our listeners' perspective, why?
Why can't I take their models and run it on my machine?
There is a financial aspect on this as well.
For example, when we are talking about closed-weight models, we can only interact with these models via website or API.
So the creators of these models, they keep the training code, they keep the weights, they keep the parameters, you know?
these are all locked in their own servers.
And these are for economical reasons as well, right?
But we are getting more and more this conversation about the AI FinOps, so to speak, because it's getting more and more expensive to use those API calls and to use those closed models from the providers like OpenAentropic.
So that's why lots of companies right now, including ours, and others out there are taking a very close look at these open-weight models to run it in-house.
Why would anybody choose a frontier model then?
Or maybe can somebody define frontier models for me?
Using the food analogy, for example, OPEX versus CAPEX, right?
So OPEX, it will be the equivalent of ordering food delivery.
So you are kind of renting access to closed modules like OpenAI, Entropic, Google via APIs.
So you pay like a tiny fraction of a cent per input, per token and output, per input and per token, essentially.
The advantage is like zero upfront cost.
So you don't need to buy expensive computers, hardwares or have a farm of Mac minis in your company or pay massive electricity bills.
You just pay for what you consume.
So pay per use, basically, right?
But with a caveat, as your application grows from hundreds of users to millions of users, your monthly bill grows linearly.
So a little spike in traffic, this translates directly to a massive recurring monthly cloud bill, you know, so that will be costly for you.
But on the other hand, if you want to use an open rate model, it's the equivalent of, okay, I'm not, I'm not.
i'm not kind of wanting to order food right anymore so i'm going to build my own commercial commercial kitchen so to speak so you make a large uh potentially one-time investment in local physical hardware so this includes you know buying desktop supercomputers uh like nvidia g g gx park or Mac Studio, Mac Mini's workstations with unified memories or even full enterprise GPU server racks, right?
The advantage on that is that once this hardware is paid for, running inference on an open-weight model like DeepSeq or GLM, this costs virtually nothing beyond basic electricity.
So your cost per token drops to near zero.
And there are also other technical advantages like the latency is lower because there will be no round trips on network or third-party servers.
Sensitive data stays entirely inside your building or in your machine.
So, yeah.
The other reason people use Frontier models, though, is that let's say Anthropic and OpenAI are some of the most familiar Frontier models that we associate with them today.
They have incredible funding.
and incredible researchers and an incredible head start.
And so those models tend to be more performant and more capable than the open weight models, which have generally smaller funding, smaller teams, sometimes unpaid teams, that kind of thing.
And so for the long time, Frontier model has directly equated to the best model.
But as we'll probably talk about over the past year with DeepSeek and the GLM series and the Quinn series, All of these models are starting to catch up rapidly to Frontier.
And so now the conversation is shifting away from good model versus bad model and more towards model that somebody runs for you and model that you can run yourself.
How close is that gap now?
Because I have to admit, as a user, so I use these things, but I don't deploy them.
I tend to want to use the most expensive model available always because, like most capitalists, I equate trust with cost.
I mean, what does that margin actually look like as far as efficacy?
I will say I am guilty of only using Fable when it's available to me and not using Opus or Sonnet or Haiku.
But what I've noticed recently, over the past few weeks, I started to play with the GLM-5.2 models, which 5.3 has made waves for being almost as good as Fable, if not as good.
to the point where there's some controversy about did they try to distill things from Anthropic and learn how they did it from the outside, which is a little, that's a little sketchy if they did that.
But either way, that's an open-weight model that competes performance to performance with Fable, potentially.
And so what I'm noticing in my own work is that the open-weight models are getting to a point where for most general tasks, like searching knowledge, summarizing documents, or writing very well-scoped and well-defined code changes, The open-weight models are just as good, at least as Opus.
Maybe as good as Fable in some cases.
What does it take to run some of these?
I mean, like, I early on tried to run some of it on my MacBook, and I have a pretty decent MacBook.
You know, 48 gigs of RAM and the big processor, the big GPU.
And most of the larger stuff I couldn't even start.
I mean, unless I didn't know what I was doing, which is equally as possible.
But what's it take to really run these?
I will say these models, the ones that are truly competitive with frontier models like Opus and Fable, we're going to have a hard time running them on a private computer.
If you have a specific kind of Mac Pro that has, say, 500 gigabytes of VRAM available to it, which they exist, but they're unavailable and unattainable now for obvious reasons.
If you have one of those, you could run these models locally.
If you have a farm of Mac minis, like Andre was saying, you can potentially run one of these with distributed inference, but it'll be a lot slower because it's having to do consensus and all the fun problems that come with a distributed system just to get you tokens.
So what I do, and probably a lot of people do, is you can run OpenWay models off of other cloud providers.
So you're still paying for your tokens.
You're still...
paying an amount of money to do inference, but you are paying a lot less upfront than having to buy a GPU, for example.
And generally, you're paying a lot less than you pay when you're using a Frontier Labs models.
Like, as an example, these aren't the actual numbers, but Fable feels like it would be like $20 per million tokens, whereas an open-weight model might be $1 by comparison.
The actual numbers, who knows?
But it's that kind of difference that you see.
By way of experiment, I was generating some Word documents earlier because that's the format I had to deliver it, the person that needed the documentation.
And I got notified by, in this case, Claude Opus that I was nearing my monthly limit, which I don't generally approach.
I have a pretty decent limit.
And I looked at it and it was charging me like $5 per doc.
It was basically what it was running.
And I dropped it to Sonnet, and of course it was a lot less.
But this is also something that I can't risk quality on.
Are there experiments that you can run?
Are there proof points?
How do you identify a good model versus bad model?
Is that even a fair thing to say?
I think, like in your case, in your example that you just said, for tasks like coding assistance, basic writing, document summarization, quick chat test, this kind of ordinary stuff.
You can do this pretty much using compact models like Lama, Quen, a distilled version of DeepSeq, like something like between 7 billion to 14 billion parameters.
So the only hardware needed, you need a standard GPU, you know.
uh anything between i don't know 10 12 to 16 gigabytes of v run like the rtx 3060 from nvidia you can also run on on the apple silicon chips as well mac studios mac minis whatever so if you have anything between like 24 32 48 gigabytes of unified memory you're good to go you know so for those type of tasks a compact model will work seamlessly, I would say.
Anecdotally, like in our engineering teams, more sophisticated engineering teams will set up evaluation harnesses for their models.
They can A, B test different kinds of model architectures, whether it's going from Claude Haiku to Claude Opus or going from a Claude family model to GPT to GLM.
And those evaluations, we could do a whole podcast on because they're a different kind of problem than conventional integration testing where there's one right answer.
Usually with LLMs and models, there's a broad spectrum of right answers, and you want to make sure that whatever it's doing falls in that spectrum.
So it's a harder challenge, but that's what some engineering teams do is set up these harnesses.
On a personal level, though, when I'm looking at do I want to use Opus for this change or Fable, or do I want to risk it on GLM, I usually look at how well the model handles uncertainty.
Because when I give a task to an experienced human engineer, and there's uncertainty in that task, they'll ask me for clarification, they'll ask me questions about it, and they'll try to minimize the assumptions they make.
What I found is that the less performant models make many assumptions, and often they're disastrous assumptions.
And the more performant models make less assumptions and usually ask for clarification.
And so if the model's asking me for more clarification and making less assumptions, I feel more confident using that model for a given family of work.
But if it just does it...
and doesn't even ask, I'm like, hmm, maybe I'll use a smarter model next time.
So, yeah.
So let's say that nowadays care is up to the user to choose or to do this kind of, you know, thinking, okay, this model is good for this, that model is good for that.
But I think that maybe we're going to achieve a time that automatically the hardness will choose, okay, this is a simple coding test I'm going to use like.
I don't know, a Lama local in your machine.
Oh, this is like I get into part of the code that I need to run a complex task here, you know, multi-step coding, deep analytical math or whatever, structured data extraction.
Okay.
It's better to use like a fable or a sonnet for that, you know?
Maybe this is a problem that's still to be solved.
It is.
And to your point, like Cloud Code right now, if you use it, it will do some of that internally.
It'll delegate your tasks to weaker models and kind of make a judgment call on could this coding be done by Haiku or Fable or Opus.
But it's still not super advanced yet.
It feels like you're rolling the dice every time you call it.
So another somewhat related topic, and this could definitely be a podcast or 12 on it, is around AI safety and and security and all of those things, the iddies.
And so I know when, gosh, gosh, it's been nine months already or 10 months or something, when Open Claw came out.
It came out right before a large event I was going to, and someone was talking about how they had done some pretty sketchy things with it, frankly.
That some of the, like...
If I had asked an anthropic model to do, it would have just flat declined.
And so, you know, what are our concerns there with people using some of these models?
I mean, someone's always going to do, there's always going to be bad actors.
So don't get me wrong.
I'm not saying that open AI is the antidote to bad actors.
But, you know, what is being done to make sure that society...
isn't being threatened by some of this stuff.
I know on the vendor side, like if you're using a Cloud Code or a Codex, one of the Frontier style, again, it goes out to the Frontier every time.
They're trying to implement tons of guardrails in their systems that when you ask it to do something dangerous, it'll decline or it'll at least be like, are you sure that you want to do that?
So as an example of how far these guardrails have come, originally when Cloud Code came out, you had to hit confirm every time the model tried to do a specific action on your computer and make a specific code change.
The idea was that you're, as an engineer, as a user, reading it, making sure it looks legit, and then saying, OK.
And then pretty quickly, people started to turn on YOLO mode, which meant that the model could just go and not ask ever, which led to a lot of people to delete their computers and do other kinds of horrible things to their systems.
And so Anthropic tried to double down on guardrails and try to make it so the model could predict when it was making bad choices sooner.
And that work has gotten to a point, at least for Cloud Code, where I think it was last week or a week before, they changed the default feature in Cloud Codes that it now accepts edits and runs autonomously without almost any clarification.
Because it's gotten to a point where they trusted enough for users that it'll do the right thing and not destroy your system.
But that's for one frontier lab.
That's not for all labs.
And so the general discourse comes back to guardrails, guardrails.
guardrails and also the lethal trifecta, which I think we've, I don't know if we talked about the podcast before, but it's the idea that, you know, a model should not simultaneously be able to take privileged action and also have access to privileged information.
The two features should be separate.
And there's a third component of it that I forget the exact definition of.
I think it was Simon Willison that coined it a few months back.
And that is a key thing is trying to separate.
functionally the concerns of how your agents interact with the world.
So they can't both read all of your emails and send random strangers DMs on Slack or Blue Sky.
Yes.
So that is potentially a hidden malicious prompt on public skills or plugins, especially with the advent of OpenClaw.
Of course, as you well said, again, malicious actors are everywhere, so they can publish...
helpful looking skills, but those skills secretly contain, you know, hidden system backdoors, credential stealers, this kind of stuff.
So we need to be very aware.
So you were talking about models earlier and you were talking about billions of parameters and, you know, it's the old joke, a billion here, a billion there.
Sooner or later, we're talking about real dollars.
What's a big model and what's a small model for some of the listeners?
Because, I mean, these all sound huge.
At least on my perspective, I would say that a small model, it's something between like 1 billion parameters to 15, 14 billion parameters, right?
So there's medium ones, you know, between 15, 16 to something like 67 billion and large models that could be hundreds of billions of parameters, right?
This is not like written in stone, but this scale is mainly categorized in those tiers and specifically based on the parameter count.
and also hardware requirements and practical use cases so for a small model if you can run in your local machine this could be considered a small model so uh that would be my take you know to easily to easily explain to a to a user uh oh i want to run on my macbook that has 48 gigabytes of ram a large model with 400 billion parameters good luck with that you're not gonna do it you know you're gonna need like a high spec farm of Mac workstations, a cloud infrastructure or Moody GPU setups in order to doing that.
And when we say large models, we are not even talking about those very large frontier models, right?
Those frontier models, they have trillions of parameters.
So they need a way more compute power in order to run.
So that's why those companies are building data centers everywhere.
Okay, because I ask because a friend that works for a bank and his boss, who was very technical, was like, well, why are we paying Cloud or OpenAI or, sorry, Anthropic or OpenAI, anything?
I read the other day you can just download DeepSeek.
And it's like, it's not quite that easy.
But, you know, when pushed back on.
His response, I guess, was, look, I have enough problem keeping up with the banking regulations that I need to do IT in my industry.
I can't keep up with what it takes to run it.
So are there rules of thumb or resources or anything you can recommend where people can say, this is reasonable for me to try to run this?
Or is every one of them really just a call?
I was talking to somebody yesterday and they showed me a clear set of breakpoints from where it makes sense to go from API-based billing to renting GPUs that you run your own models on to buying your own GPUs.
And there's different schools of thought on it.
But I want to say the numbers, once you reach past, I think it was like 4,000 tokens per minute or something, that part I'm not sure on, but there is a breakpoint.
at which point you should start buying your own infrastructure and running it on there.
That said, like any service that you decide to build yourself instead of buying from somebody else, you're now inheriting all the operational burden that comes with it.
So even if it makes financial sense, purely financially, to rent a GPU versus pay for tokens, it might cost you more in terms of ownership and operations to actually operate that service correctly such that it still works for your engineering teams.
Because that is a very different part of the question entirely.
Sounds a lot like cloud not that many years ago, where when it first started, it was all the same CPUs.
And then, oh, it's specialty.
And now, closest data center and what have you.
Is all this shaking out?
I mean, are the big vendors responding and trying to find alternatives?
Or are they just saying, too bad, pay me?
One of the interesting arguments I heard was that big vendors are coming under a little bit of, not like fire, but heavy competition from the smaller neoclouds.
Because you have these incumbents like Google Cloud or Azure or AWS who have these massive data centers, but their data centers are heterogeneous.
They've got all kinds of different hardware and they're doing all kinds of different things for different tasks.
Only some of that is hardware that has GPUs that can run models.
Whereas the neoclouds are coming in and saying, hey, all we care about is serving models at scale.
And so they're building more specialized data centers that are just GPU after GPU after GPU.
And so their unit economics are in favor of running models so they can undercut these larger, more established cloud providers.
At least that's the way the thinking goes.
I don't know how that's playing out in reality yet.
I personally think that there are lots of hardware makers entering this space.
And with the intent of intentionally redesign the developer setup for this local execution, right?
For example, I was reading on TechCrunch, I believe this week or the week before that the new wave of Apple computers, they are going to be prepared for running AI locally, you know, because Apple developed those silicon chips and whatever.
So...
And also, NVIDIA, for example, NVIDIA launched last year the DJI Spark, so, which is, it's still expensive, it's around $5,000, if I'm not mistaken, for the little box with a GPU inside.
But it's bringing, you know, desktop class data center to your house, you know, it's bringing power straight to your desk for local fine-turning, heavy prototyping.
But the other thing that gets me more excited is the advent of physical AI.
So we have vendors now developing hardware.
For example, I believe this week or last week, Arduino launched the Arduino Ventuno Q, which is powered by a Qualcomm chip design.
So imagine have this kind of edge hardware.
in your table, in your desk for you to play around, you know?
And this piece of hardware is capable of running quantized LLMs, VLMs, you know, directly on a little machine and especially on robotics and industrial devices.
So this physical, I think, is one of the things that excites me the most, I would say.
So you were talking about hardware.
You did a little bit in our prep when you mentioned a few different pieces of hardware.
I know some gamers like we all do, and they wanted to go out and get like a 5090 or whatever the card is these days.
And I have an older, older, 4070, which is not old really.
But it's only 12 gigs of VRAM.
And what it came down to is they just couldn't find one.
And when they could, the street price was double what NVIDIA is.
So like you just mentioned an NVIDIA offering.
You said it was 5,000 US.
I'll bet you can't get one for that price in the U.S.
I'm sure people are marking them up.
Are you seeing any light at the end of the tunnel for hardware availability for this?
Because, I mean, it's great if we can do it, but if there's no fuel for the car.
I have hopes and dreams.
There was a split, I suppose it was last year or the year before, but a major researcher, I believe it was from Meta, left.
to start a startup working on world models.
World models are this idea that instead of having an AI agent that's trained on text and predicting text and doing all this stuff, it understands the actual world that it's operating in.
So like our physical space are the same way that we as humans kind of understand our world.
And these models can solve a wide range of tasks, presumably also text generation tasks, because that's part of our world.
And I just hope.
And as part of that line of research, we might find an architecture that's not as GPU heavy in the first place.
And that's more what I'm hoping for is that advances in AI will push us away from these billions of gradient descent trained weights that we rely on and towards a slightly more intelligent architecture that doesn't need 500 gigs of VRAM to operate.
But that's a hope.
Also, on the other side, we hear that...
these memory makers, they are currently building new fabrication plants, new facilities to meet these AI memory demand.
And we hope that all these facilities come online, come up to the live.
So maybe next year, who knows?
So the total global supply will surge.
So this extra capacity will relieve the pressure on consumer.
on the consumer of memory, GPUs, et cetera, et cetera.
Andre, you said a little while ago that you reduced it to the cost of electricity, which is certainly a lot cheaper than these vendors are.
But I know that, like, I was in London a little while ago during the heat wave, and electricity, it's not just the cost, it's availability.
You know, most of the air conditioner in my hotel didn't work.
And it was 102 Fahrenheit, which I think is, I think the official measurement in Celsius is hot.
But, I mean, it was pretty bad.
And it was because there wasn't enough power.
It wasn't about the cost of the power.
There just wasn't enough.
So do you have any concerns there with people standing up their own data centers?
Yes, like electricity.
The energy problem is an important subject, right?
the whole population of the earth firing up tens of thousands of active parameters across those server racks, you know, they are all drawing hundreds of watts per credit, you know, so this is a thing that we need to be aware.
But I think that one way of solving this is kind of using...
I come from the infrastructure.
I have a...
infrastructure background.
So I think that one way of solving this is like having different deployment strategies to shift away from centralized cloud data centers.
Right.
So again, run those models directly in our hardware, our laptops, our Mac studios, our phone, our local wide servers.
Right.
Because those local devices, they draw a fraction of the power, right, of a data center.
So this could be one thing.
Also, another thing is that like, which like at least here in Brazil, people are doing a lot, bringing your own power, you know?
So people here are increasing the coupling from public grids by using solar panels.
So there is a lot of houses using solar panels here in my region.
So I believe we can use this strategy, so to speak, to diminish our electricity bill and to save the planet.
I think one of the most staggering facts I learned, I think it was last year, and I think it was Microsoft, I'm not sure, but one of the major cloud providers has signed some contracts to bring nuclear power plants back online, specifically to power their AI data centers.
And I was like, well...
Nuclear power, if it doesn't melt down, is probably cleaner.
But also, that's crazy that that's the point we've reached, is nuclear power plants to power data centers.
So I think that the major clouds are finding creative ways to get the energy.
But to Andre's point, I think the real push needs to be more towards edge computing.
And we started to see that as a trend in systems engineering in general over the past few years, or more and more.
efforts in putting the decentralized systems towards pushing compute to the edge, all that, which is a good thing.
I love that.
And I think that'll be a key part of how I respond to what will probably be an energy crisis from AI at some point, if it doesn't stop.
Yeah, we don't need another Chernobyl.
Yeah, you know, it's funny.
I was talking to our chief AI officer, and he was looking at some of that.
And I don't remember the technical term, and I don't want to misquote.
But he found it very promising.
some of the little local things.
And I actually spent a bit of time on nuclear-powered ships when I was not quite this gray.
And so I admit I'm torn because, you know, I mean, I don't really want to be like Chernobyl.
But we've been running small reactors pretty successfully for a very long time.
And so it's like, I don't know.
It's scary.
Anyhow, I digress on purpose somewhat.
So I guess where do y'all see this going?
Like, is this a flash in the pan because OpenClaw want everybody to have their own thing?
Is this no finances are going to do it?
And like, I'll have to admit that.
So I do this technology podcast, but I also do some of our business facing things.
And it's been very rare.
You both have mentioned either CapEx or OpEx during this podcast recording.
For those that aren't familiar with those terms, it's the capital expense, a thing I buy, versus an operating expense, a person I pay or a thing, a bill or whatever.
And that hasn't creeped into technology conversations since cloud 15 or 16 years ago.
Are you now hearing these terms from your technology clients?
So most of my work at ThoughtWorks is admittedly internal.
So my clients are ThoughtWorks and ThoughtWorks leadership.
But concretely, the CapEx versus OpEx conversation has shifted with this frontier models, especially because we saw just a month ago when Fable got shut down for a few weeks because it was a little bit too dangerous.
And that's not an OpEx cost issue, but it highlights the fact that if you had depended on Fable and it was an external resource that you didn't control, you were suddenly in a lot of trouble.
And so that's sort of an operating risk that you're carrying by using a frontier model that you have no way to run yourself.
But then in addition to that, prices for models have been kind of going up and down a little bit and are a little unpredictable.
And so the OPEX for these frontier models is very unpredictable and kind of scary if you're a technology leader saying, I'm going to build this platform and it's going to cost this much.
You can't say that because you don't know what it's going to cost in six months.
Whereas at the CapEx conversation of I'm going to buy a GPU and run this model on it, that's an expense you pay once and you know what it is and you can amortize it and it's great.
So it's a very different conversation.
And what helps the conversation is that the gap between frontier and open-wind models is closing.
Frontier models are getting smarter and smarter, but there is a certain point of smartness where it's good enough.
And the open-weight models are really starting to cross that, especially with the GLM-5.3 and the QIN-3.8s.
Yeah, the industry is moving away from the paradigm of let's only use those larger models hosted in giant cloud data centers.
It's moving away from that toward efficient, localized, specialized data execution.
So as you perfectly pointed, Ken, It's kind of the same thought of cloud computing in 2015, right?
And I always joke that our area, our field, IT slash software, it runs in circles.
So it's that time again, right?
So I believe the macro trends to shape this next era of AI is like using frontier models in the cloud for ultra heavy reasoning.
or base model training.
And for the majority of tasks, we are going to run specialized small language models, you know, for domain specific tasks, because those have fixed costs.
The latency is better.
And also companies are doubling down on local hardware execution, you know, as we just said before, like running on your own device or on micro data center.
So We can say that the right sizing replaces this scale, this giant scale.
We could keep this going for a very long time, but for the sake of our listeners' sanity, what can people do tomorrow when they go to work?
Friday as we're recording this, so Monday morning.
What's actionable?
I mean, is it just run experiments?
Is it read more?
What's an action that people can take when they get back to work?
And I guess, Andre, you started, so this time I'll start with care.
For me, if somebody is in an engineering function, and so they're writing code every day and they're using agents to do it, I would encourage them to experiment with using an open-weight model as an alternative to whatever frontier model they might be using today.
Especially with the fact that GLM 5.3 has become widely available this week, that one is a great one to try first, because I think you'll find that it works almost as good as or better than Opus.
for a lot of tasks.
And that was the premium, premier model for a while now.
And Andre?
Yeah, I agree with Kerr.
So I would add, like, don't wait for a formal enterprise budget approval, you know, in order to use your experiments.
So download a local model, local runner, like OLAMA, VLLM, grab a terminal harness, you know, like open code, pull up an open wait coding model and...
And do your thing.
And don't spend $5 per document.
That's the most important thing.
But the compliance, Andre.
The compliance.
Yes.
I wish this was a vlog, a video blog, and not an audio podcast because care has an internal role and I work for the office of the CTO and I agree with Andre, but both of our faces went, oh, no.
If it was that easy, right?
Yeah, sometimes it's not that easy.
I say, yeah.
Just please be careful if you do that.
I heard about many executives running open claw over winter break without asking their IT teams and getting all their inboxes deleted.
So I guess that's a warning.
Be careful.
And on that note, thank you, Cara.
Thank you, Andre, for your participation.
Thank you, as always, for the listeners.
And we look forward to hearing more from you soon.
My pleasure.
Thanks for having me.
Thank you.
