# Open Source AI Enterprise Adoption Strategy

**Podcast:** Y Combinator Startup Podcast
**Published:** 2026-09-12

## Transcript

Cost is by far the largest pain point that open models can jump in and solve.
But, you know, every business has a vision of getting better control over AI and customizing it for their business.
And that's really their North Star.
You know, cost is something they can solve in the short term, but it then enables them to then go and customize these models for their unique use case.
Early 2024, there was lots of interest in.
fine-tuning your own custom models then it sort of went away or that will just be wasted effort or get stomped by the next model release seems like it's coming back now you have a front seat to all of it do you think we're going through like another cycle or is it here to stay this time Welcome back to another episode of The Light Cone.
Today, we're talking to Jeffrey Morgan, co-founder and CEO of Olama, the easiest way to run open source AI models locally and in the cloud.
Olama is used by 9 million developers, has 178,000 GitHub stars, and is used by 85% of the Fortune 500, which means Jeff knows a lot about the state of the art of AI, what models score highest on benchmarks, and what developers actually download and keep using.
Jeff, welcome to the light cone.
Thank you for having me.
We're down to here.
What is the state of the art?
What are you seeing out there?
Well, I think the biggest thing we're seeing is a shift to open models, especially in enterprise.
And that's from a mix of US and Chinese origin models.
And it's predominantly driven by coding agents and also AI assistants, more co-work use cases like OpenClaw and Hermes.
And because you sit in the token flow of like so many tokens, you have really good data on what models people are actually using and how it's changing.
What are the trends that you're seeing?
Yeah, you know, Olamo started as a way to run open models on your MacBook or other hardware, NVIDIA, AMD, Intel.
Earlier this year, we launched Ollama's cloud.
And what we're seeing there is that it's predominantly Chinese models right now of Chinese origin, but they're being accessed by businesses all over the world, especially US and Germany is actually a big source of where open model tokens are being accessed.
Is it all about cost?
Is it our enterprise is coming because they just want to get the cost down or is there anything more?
to it?
Cost is by far the largest pain point that open models can jump in and solve.
But every business has a vision of getting better control over AI and customizing it for their business.
And that's really their North Star.
Cost is something they can solve in the short term, but it then enables them to then go and customize these models for their unique use case.
Is there a particular large enterprise that you can name that has done this?
I think there was a great article in the information yesterday from AT&T, and it ends up they've already shifted 40% of their token consumption to open models.
And that's right now predominantly through U.S.
and Europe models, but they're also evaluating the Chinese models.
What kind of workflows do they run?
Predominantly coding agents.
I think what we've seen just from the extreme growth and...
you know, per developer or per user token usage has predominantly been from coding agents.
And then earlier in March and April, we saw OpenClaw take off and subsequently the Hermes project, Hermes agent project take off, which has then opened up that ability to automate a huge chunk of work over a long span of time to non-developers too, whether it's like finance or support or marketing or sales.
You had this actually very cool graph on the takeoff.
exponential for OpenClaw.
Yeah.
So earlier this year, this is a graph of token usage by developer on Ollama's cloud, the average amount of tokens they use per week.
And we kind of start- This is per developer.
So this graph looks like it should be an aggregate of Ollama's growth, but this is actually the per user.
Exactly.
This is on an individual user basis.
How many tokens are they using a week?
And so there's kind of like two big inflection points.
One is that initial run up at the start of the year, which was driven by coding agents.
So we saw Kimi, the GLM models, Minimax launch.
Finally, we had open models that could power coding agents.
And then in April, we saw this incredible growth from...
OpenClaw, really, which was then not just developers, but the rest of the world could take a hard problem, give it to an open model and let it go complete the task, which obviously consumes a ton of tokens as it's figuring out what tools to use, what data to go fetch.
We went from a context window of 128K to a million plus with open models.
And so all that enabled this explosive growth.
So it went from roughly 5X, so under somewhere around, I don't know, 15 million tokens before.
before all these co-work type of use cases?
I think that's about right for that open claw jump we saw in April.
Obviously, in aggregate, it's in the 10 to 20x, if not more.
As a whole, through Ollamas Cloud, we saw 150x since the start of the year.
And so it's just the big interesting thing here is this.
huge surge in demand for open models, right?
Whereas open models, I'd say in 2024, 2025, from the large models being served were mostly being served as custom models.
So you'd taken off the shelf model like DeepSeek or Kimi, and then you'd fine tune it for your use case.
Like for example, you know, Cursor had famously done.
And from there, you know, you could serve that at scale.
But seeing out of the box open models being served, that really only took off at the start of this year.
Even since we started this podcast, these things sort of come in.
cycles maybe like i feel like early 2024 there was lots of interest in fine-tuning your own custom models then it sort of went away and it's like or that will just be wasted effort or get stomped by the next model release seems like it's coming back now you have a front seat to all of it like what do you think we're going through like another cycle or is it here to stay this time i think the release of these models the cadence is only speeding up which makes it ever more harder to you know stay on top of that and have a post You mean the latest closed frontier models are releasing faster than ever?
I think on the open source side, it's getting faster and faster.
You know, this summer, we've already seen three iterations of the DeepSeq Flash model as an example.
What used to be more of a six-month cycle.
Yeah, so the gap is like closing, essentially.
It's closing.
And I think that makes it even harder to custom train models.
On the flip side, I think the tooling is getting better.
And so it allows teams that want to fine-tune their models to stay on top of it.
I mean, we're also entering this moment where AI safety is becoming more and more of an issue at the Frontier Labs.
So the Frontier may well slow down to figure out its alignment and containment issues.
And then meanwhile, the open source models and open weight models are continuing to grow and get better.
Yeah.
And, you know, we saw the announcement and the release of the GLM-5.3 model and its capabilities from a cybersecurity standpoint, you know, being extremely impressive.
It creates a big opportunity for whether it's startups or existing businesses in the security and governance space to really jump in and help.
Because I think if you look at, you know, that AT&T article I was speaking about, it largely, the blocker to adopting open models is largely around security and safety.
But from our experience talking to customers, whether it's in Europe, whether it's here in the U.S., if you can solve the safety problems, by and large, adopting the Chinese origin model labs is completely on the table.
And it's really exciting for these businesses.
There's sort of this interesting moment right now where a hugging face had to use open weight models to actually even detect the hack from the frontier.
A common question we get is like, well.
where can I use open models that are leaps and bounds of an advantage over using a frontier closed model?
And one of the key use cases is security testing and making sure that your software is secure.
Yeah, because if you try to get Claude to pen test your product, it will just refuse to do that.
Correct, yeah, by and large.
Whereas there are literally obliterated security researcher models that you can find on Hugging Face that allow you to do it.
That is true.
There are ones that are custom trained to be even more liberal to go and attack these problems.
But even the out-of-the-box models, they do come with...
safety training, but they're a little better understanding if you're doing this for a good use case versus one that's more of a negative or malicious use case.
A cool thing about Olama is that because you guys are such a key distribution channel for these models, my understanding is that...
typically the model developers are contacting you before the general release to coordinate launches and stuff like that.
And you often get sort of like previews of what's about to happen.
And you get these incredible growth spikes when it...
new model drops.
Like, I wonder if you could like tell us a bit about like what it's like to operate this thing at scale.
Yeah, absolutely.
And like I said before, the models are coming out faster and faster and faster.
And so we've developed a playbook to what does a successful day zero model launch look like?
And there are a lot of things to get right.
There's making sure that it's supported and your favorite inference engine, which is generally sometimes a multi-week process to make sure it's fast, to make sure it's accurate, to make sure it's up to spec with the reference.
There's also finding the right use cases and harnesses for developers and users to make use of this model.
Generally, these models have new capabilities.
This morning, DeepSeek launched their first multimodal model.
From a large model LLM standpoint, they had previously had some smaller OCR models, which unlocks a whole bunch of use cases.
But what they also changed was the DeepSeek harness, which launched recently so that it could support this capability.
So step two is then to find the harnesses, make sure they're prepared to actually run this model and to do that effectively.
But every model is different.
They all have different challenges and architecture changes and tool calling mechanics.
And getting all this right is really hard.
The key thing to do at the end of the day is to run benchmarks against the final product ahead of release and make sure that it's running as the research team at the lab specified.
So I think one very important part that you play in the whole ecosystem is you kind of create a very legible standard to be able to...
Make sure you have the best way to use the harness for each new model, each new tool call, and all of it is consistent across all the different models, which is pretty hard to do.
Yeah, I think there are three things really that, you know, we try to package together.
One is harness, for example, you know, whether it's an off-the-shelf harness or an SDK to help use the model, some existing harness that is designed for this.
And what's great now, there's so many great.
open source harnesses the codex harness is open source uh open code's a great one and one of the most popular harnesses from olami users step one's getting that right but then you got to package it with the model making sure that it's you know available it's reliable if it's in the cloud that there's enough capacity for it because day zero tends to be the the largest growth day, obviously.
And then importantly, there's the hardware and the providers.
And that's actually where there's a lot of collaboration to be had, whether it's like an inference provider optimizing the model or, you know, with some of our partners, whether it's NVIDIA or, you know, for example, working with the Apple Silicon stack, making sure that it actually runs the model really fast because of.
the model's capable, but it's really slow.
That's not a great experience.
So getting those three things packaged into a box, and generally you get the model, you're lucky if there's the ability to access the model a few weeks in advance.
A lot of this stuff comes together in the last 24 hours before the model gets released.
And so it's generally a fire drill.
So you kind of have almost like an operating system in the old world where you needed to really integrate very tightly with all the drivers, all the hardware, and then at the application level to make sure that all the apps were...
really tuned up well and you are the glue for all of it right i do think the os which is generally a cliche analogy to use is a good one because you've got the drivers for the hardware and the providers and the inference layer but you also have the application runtime and making sure that the harness works gluing that together it's a very combinatorically large problem to solve and so doing that well is really difficult and so but you know over time you develop pieces that you know allow you to quickly develop that, test it, release it, and kind of have this common runtime that can match any harness to any model.
And that's kind of the role we're really playing for developers.
You guys are very hardcore engineers and you fine tune things all the way to, from the origins with Apple Silicon, all the way now to DJX.
How do you build such a deep technical bench with that?
I think a lot of our team, you know, we aren't AI researchers by background.
We're from VMware and Docker and from, you know, other networking companies.
And so the classic compute problems are kind of reinventing themselves in the inference land, whether, and so largely, you know, that's where we like to focus our time.
But I think at the end of the day, you know, it all comes down to the developer experience.
what is it like when the developer makes a call to the API and gets tokens back?
What happens?
And there's more and more happening in that layer right now.
And getting that right needs all the layers of the stack to work well together.
And I think, you know, one of the challenges with open models has been that hasn't been happening at the rate of what a frontier model lab puts out where they have, you know, the classic five layer cake, right?
That Jensen mentioned, which is like the apps, you have, you know, the model, you have the infrastructure and inference, you've got the chips and then you got the energy and they've got all that ready to go for developers on day zero.
And that's really the thing that we're trying to reproduce for open models.
Of course, we're not going to do every layer of the stack, but we can help orchestrate that.
You layer the layers.
Yeah.
And maybe in open models, there's more than five layers.
Like that model layer actually has a lot to it, right?
There's the model weights, but there's also a lot of the orchestration components that each model is uniquely good at.
There's this developer API layer.
There's so many opportunities to build.
between the model and the application layer that are kind of hidden in today's five-layer cake stack.
That's interesting.
Do you think any of those hidden layers might get unbundled and become their own companies?
providers?
Absolutely.
There was a really good talk from some of the Anthropic team, the platform team, and they talked about three big things.
One was knowledge.
How do you connect your company's data and context to the model?
One is coordination.
As you know, when you make a request to in your Claude app or Codex, it goes off and spins out a bunch of sub agents, some of those in the cloud, some locally.
There's a coordination problem.
And then lastly, there's an execution problem, which we talk a lot about as sandboxes, but These agents, more and more of them are moving to the cloud.
There's a huge compute problem to be solved there.
And if you look at what's happened classically in the cloud business, open source has meant that there are best of breed companies for each of those problems.
Whereas you might have had kind of the, if you look back to the original generation cloud products, you have like the Heroku's of the world, you have Google App Engine, where all those things were bundled together.
But what developers ended up preferring is best of breed products for each one.
It's the novalism, which is all things are just bundling or unbundling.
Exactly.
Like the frontier model.
Frontier model labs want you to be totally in their walled garden of managed agents and their context, their memory layer.
And then meanwhile, Little Tech and all the founders out there and all the open source developers don't want to be caged in.
So we're going to make all this other stuff and it'll be...
an interesting moment to figure out like what ends up winning.
I mean, it'll probably be some mix of both.
I think so.
And you know, you've got this abundance of open model tokens that's being created.
There are just dozens of open model providers that are able to serve these tokens.
The new scarcity, the problems now are what's above the tokens, right?
You know, how do you orchestrate an agent from, you know, A to B?
These are problems that are tons of new you know, systems and engineering problems that are just really hard to solve for an individual dev.
Like there's no way they're going to build all those layers.
Well, the interesting thing now is because the coding agents themselves are getting a lot better and ostensibly this is the worst the models will ever be.
The classic reason why there was a moat here was it was just too hard to have really well-maintained software that was properly tested that actually satisfied user need.
And what if that goes away?
Like we're literally at this moment where Actually, maybe the, you know, you'll just have a cron and it runs a markdown file on some TypeScript and it'll just, you know, there's no lock-in anymore, right?
Like you could be using OpenAI's memory system one day and then actually, you know, you could have an agent be constantly syncing that against your own memory system and actually it just works.
It's fine.
Like it works over MCP, like there's all, you know, there's plenty of like runtime testing and then there's no lock-in.
Yeah, I think a lot of the, you know, what we know of as harnesses today, a lot of those pieces will go down into the model.
But, you know, as you kind of push a lot of the core loop of the model and hooks down to the model itself, there are these pieces that kind of come out.
Like memory is a great one.
And the general guidance we're using is if it's a stateful problem, like there's storage involved, that's something that, you know, in the end.
can't go into the model because the model's trained and it's, you know, as we know, training runs are now happening on like a monthly basis, but it's still not up to date with the latest data.
So generally storing data is like a huge problem space that I don't think will ever make its way down into the model layer.
Like managing credentials and that kind of stuff.
Security, credentials, safety.
Open models don't have all the safety tooling that closed model providers give you out of the box, but that's super important, especially for businesses to adopt them.
I'm curious where you see sort of the the end or future state for enterprise on the balance between sort of frontier closed models and open source model um especially on spend like my it feels like you know initially it was just the everyone's just like allocating all of their budget to anthropic or um open ai uh my sense now is yes open a open source is clearly growing but like so are so is like anthropic spend so like the things seem to be growing together like does that continue or do you think there's like a steady state where it's like i don't know like half the budget's going to be on the closed source frontier model and half's going to be an open source or something different the super majority of tokens and this is our take it will be open models within a business call it 80 90 that doesn't mean 89 of the budget will go to the models in fact i think what the open model community is doing incredibly well together is lowering the cost to make it more accessible and so maybe you'll only pay 10 to 20 percent uh of the cost towards open models but your most of your tokens will be going through the open models right which will enable a whole bunch of use cases on top because you have this abundance of tokens you're not thinking about taking away token access from your team you're giving more and more access i think for the hardest tasks that's reserved for these frontier labs where a lot of the best researchers are and then from there there's a whole bunch of problems in the middle right where maybe it's a combination of open and closed models working together i think the steady state is that most of the software's most of the models are open It's kind of an interesting idea because as the models get more powerful, ideally you just want to delegate to like your smartest model to figure out when to go to like an open model.
But the labs who are in the models presumably don't want that.
And I think, look, I think that all the labs are aligned in many ways to one thing, which is how do you serve the customer?
And I think it will be up to the customer to decide if I have a router where some of the scheduling and harder, you know, orchestration happens through a frontier model, so be it.
But a lot of the kind of line item work can happen through open models and the collaboration of the two together.
And I think we've seen a ton of projects, whether it's from Sakana AI or open router that have combined the two and it seemed really good results.
This is not too dissimilar from a human organization, right?
Like you have, you know, like a law firm, there's like a partner and then there's like a bunch of associates and like the partner farms out the work to the associates, right?
It's like the same.
We saw the same thing with cloud computing where it was really a blend of proprietary software, some of them provided by the cloud providers themselves.
For example, you know, AWS had DynamoDB, which was kind of their proprietary scale-out database.
But then a lot of customers use that in conjunction with PostgresDB.
And in the end, you know, what we see as customers will use a combination of the two.
I think it's a very common pattern.
I mean, this is also the same design for why the Apple Silicon is actually more superior, the special accelerators for different kinds of workloads for, let's say, image processing, as opposed to audio.
That's been, or even go way back in the PC era, you had like your standalone audio card, right?
Graphics card and all that.
Speaking of Apple Silicon, should we talk about local models?
Because you're in a bit of a unique position because you have large businesses, both in cloud-hosted models and locally hosted models that will run on your laptop.
What are you seeing in those two worlds and what do you think is going to happen?
I think it's incredibly exciting because it's similar to the closed versus open model question.
It'll be a mix in our mind.
And that's what we hear from customers as well, where for easier tasks, you could run them locally and with lower latency and of course, lower costs when it comes to the per token costs.
Ultimately, you're...
buying hardware up front.
And you'll use that in conjunction with these cloud models.
What's exciting about this next generation of hardware, which we've had for a few years now, is just how good they are at running the 20 billion parameter to 40 billion parameter range of models, sometimes up to 120 billion parameters.
Yeah, Quint 3.8.
38B is now as good as Opus 4.6 for coding.
Is that right?
That's what the benchmarks show.
Yeah, that's wild.
It's incredibly exciting because you can run that on not the lowest memory MacBook, but the second lowest memory MacBook you can buy from the store.
So it's incredible.
And are you seeing your users do that?
How are you seeing people use the local models versus the cloud hosted models for in practice?
From the side of...
which models they're running, we're seeing a really solid mix of US and Chinese trained models being used for local.
And we have the incredible models from the original Lama models, of course, but also the Gemma models from DeepMind.
These are great choices for local.
But when it comes down to use cases, coding agents by and large are most effective with the large cloud models.
You're solving really hard problems.
You're writing code, tasks.
It's really difficult.
versus some of the document processing workflow use cases that run extremely well locally because they don't have as difficult of a task in the end-to-end, you know, problem you're trying to solve.
And so that's where we see this hybrid execution model where some of the easier, more straightforward tasks run locally.
And then you have a router that can help decide, hey, we need to go to a large cloud model for this.
And I think what that means for customers is that you're really dropping the costs even further.
when you're going from open models, not just because they're cheaper to run in the cloud, but because now you can run them effectively for free on the hardware you're buying for your business anyways.
And what we're seeing ultimately from the cloud coding agent models is it is predominantly Chinese models being consumed today.
And for the local models, it's a really strong blend of US, Europe, and Chinese origin models.
Yeah, these two graphs are pretty stunning in comparison.
Basically, for local models, the US and Chinese models are neck and neck.
We're tied.
And for cloud-hosted models, the US is recoloring the x-axis.
It's 100% Chinese models.
Basically, we need more US labs to make large models.
Is that what this graph is showing?
Effectively.
And with the launch of the NemoTron Ultra model, we're seeing...
kind of the first wave of that and it's really exciting nvidia as a company is so interesting because um their moat is not like trying to start new software businesses or sell you know tokens they seem to be quite interested in just releasing a lot of open source and helping the ecosystem and then the fact that they do that then helps them stay ahead of the game on the hardware side I think so.
And ultimately NVIDIA, what's so incredible is there's helping power an ecosystem around open models, whether that's the hardware, the models.
We've seen the new DGX station computers that they're working on, which have a GB300 on your desk.
How do we get on that list?
That isn't deafening loud.
Do you know the price point on that thing yet?
I don't know it off the bat.
Gary's buying it.
Well, I looked it up.
I mean, you can probably run a frontier model for like, I mean, very slowly for like two, $300,000.
Is that, is that right?
I think it's even more competitive than that.
Yeah.
And you can run more than a frontier model at high speeds at a price point that isn't very.
far off what you can buy from a classic workstation computer.
Oh, no way.
If you think about, you know, quite a few of the customers we talked to, some of them are banks, for example, or industrial businesses.
They already have these NVIDIA workstation GPUs in every single engineer's desk.
Some of them, tens of thousands of them.
And so this is.
Oh, I want one.
They've been selling them for all the CAD work.
Well, this is for, this is for the kind of original RTX A6000s, but this is the next generation.
Yeah.
I'm going to have to email Jensen.
This was at GTC.
So what we see here was at GTC.
And we were one of the first people, along with Elon and a few others, to receive the DTX Spark as well.
which sits on your desk and provides 128 gigabytes of unified memory to run that kind of 20 to 120b model range.
But obviously that's just the beginning of a whole new range of hardware that can run the biggest models.
So did you say you can buy a bunch of these and chain them and actually run a 400b model?
You can, absolutely.
They have this really fast network link.
And so you can stack them on your desk, almost like a miniature data center rack.
The thing that people do with the Mac minis?
Absolutely.
Now this is like the production version of it.
I think that's what's so exciting is you're seeing both from Apple and NVIDIA, this incredibly, incredible leap to next generation hardware that's built for these models and is effective at running that.
So you'd say like, this is, this is the platform to get, like you could make Apple studios work, but like if you want something that just.
can work, get DGX Spark.
From our testing, both are very competitive.
Got it.
So I think a lot of it will come down to what you can get.
And then also, the tech stack you're looking for.
I think there's an incredibly mature tech stack through the MLX project with Apple, where they've done some amazing work to run LLMs on the Mac Studio, but also the smaller Macs.
And of course, the DGX Spark stack is just incredible.
We're super excited as partners with NVIDIA for that.
And it's going to be a cool renaissance for personal desktops.
I think so.
And you know, it's funny with Olama's journey.
We started local.
Clearly the coding agent demand is in the cloud, but that's going to come back locally in our minds because the hardware will catch up when you have a GB300 on your desk and you want the fastest coding loop.
That's as fast as running your tests or as fast as making code editor change.
We all remember the GitHub Copile experience of having the autocomplete come up in a few milliseconds, a hundred milliseconds.
That experience will make its way back to the desk, which has just been a journey of starting local, going to the cloud.
And then we think that'll come back local and you'll end up using the two together.
Speaking of like what you can get in order to run.
a llama cloud, you need like a shit ton of GPUs.
What are you seeing in the GPU market?
I think what we're seeing is ultimately the prices are changing very quickly and the supply and demand volatility is very high there.
And so, you know, I think ultimately for if you're a startup, getting access to some of the B200, B300 GPUs you need to run these latest models is very hard.
Thankfully, there's a great set of inference providers building on top of that.
And so we're seeing this extreme demand.
Are you able to get all the GPUs that you need?
Are you constantly growth limited by how many GPUs you can get your hands on?
What's the current state?
We're lucky in that we've partnered with quite a few providers to work together to pool a bunch of GPUs together, which allows us to stay on top of our demand.
But that's a lot of work.
And it's definitely a lot of spending time thinking through, you know, which model will get run where, how fast should it be, which region is it in, what will the latency be for the customer.
There's a lot of hard problems to solve in that stack.
And I think what's really exciting about products like Open Router, Olama, the Open Code project is for an end user developer, they can sign up and get access to this without having to go negotiate prices on a B200, B300.
you know, think about their 24 month forecast in order to get access to some of these these, you know, GPUs.
YC's next batch is now taking applications.
Got a startup in you?
Apply at YCombinator.com slash apply.
It's never too early and filling out the app will level up your idea.
OK, back to the video.
Suppose you were like a.
a startup founder and you are just starting out now and you're building some AI company and you haven't raised a lot of money.
And so you like want to like use as many tokens as possible, like inexpensively, like what would your advice be to that person about like how they can get like huge mileage with like a limited budget?
There's this new class of models like DeepSeek Flash is a great example.
And I think there'll be quite a few more where it's ultra low cost per token.
It's also low cost per task, which is a really important metric.
And that class of models...
in my mind will be the first ones that come down to this idea of like unlimited tokens.
We all remember ChowGBT, you didn't really have to think about how many tokens you were using.
You would just use it every day, you had unlimited.
Ultimately, I think we return to that, but it's gonna take a lot of work in the model, the architecture to be custom trained for high volume token usage.
And if you think about the start of the year, we really want open models got to the frontier of intelligence.
We've bridged the gap.
We're maybe like less than three months behind between the frontier closed models and the open models.
But the next problem to solve is extreme efficiency.
Seeing, for example, the GBT Luna model become very, very price effective for customers has been a huge boom.
We talked a ton of customers where that kind of pricing enables widespread adoption within a team.
I think we're going to see that with open models.
We already are seeing that with open models.
I think the DeepSeq Flash model is leading that charge.
If we go back to the model breakdown on Alamas Cloud, the highest growth area is definitely the DeepSeq model.
And this is largely powered by the DeepSeq Flash adoption.
So this new class of flash models where they're good enough for 80% of the tasks, they're really fast and they're ultra cheap.
There's new class of model that I think will enable some of those use cases.
Yeah.
Those are going to be like the workhorse models to do like all the grunt work.
Exactly.
Yeah.
You won't have to be thinking about how many requests am I making, how many tokens you'll be much more inclined to consume as much as you can because, you know, it's able to solve the hardest, not the hardest, but, you know, difficult problems.
If you think back to the coordination layer we were talking about earlier too, being able to coordinate these flash models together to do different tasks can also yield great results that a bigger model can.
So by having these cheaper models, not only are they more accessible, they can run faster and you can access them in higher volume, but you can start to chain them together and build new problems.
that are solved by orchestration on top.
And that's a really exciting area for new startups, for existing inference providers, for some of the larger businesses today that solve workflow problems.
Ultimately, being able to chain these models together is going to be super helpful.
And you won't have to think about the underlying costs.
Yeah, I guess, you know, when we first started talking about AGI, even on this podcast, there was sort of debate about, you know, and I think a lot of AI researchers would come out and say, like, there's just going to be a giant God model and it's going to do everything.
But, you know.
I think so far, like it hasn't quite worked out that way.
Like obviously you still have, you know, if you have to literally hack the NSA, maybe you need mythos or something.
But for the majority of use cases, like you're talking about orchestration and you're talking about like smaller models, you know, that the task composition actually probably gives you a bunch of ways to make it more repeatable.
It's more trustworthy.
Like it actually does work at a cost that is...
like possible so you know if it was going to be um god model versus like lots of you know smaller special purpose or even just like simpler models um it's turning out to be the latter so far i think for most customer use cases there's a you know level at which a model becomes good enough and then they can continue using that level of intelligence maybe the model will get faster it'll have better architecture it'll be have new capabilities but they won't have to reach for the God model.
But I do think there are use cases where the most powerful models unlock them and that'll continue to be a thing.
It'll be really exciting.
You know, sometimes scary as well on what they can do.
But for the run-of-the-mill use cases where open models really shine, I think that's where, you know, we're hitting a point where, you know, you're not solving necessarily the hardest problems within the business, but they're hard enough where it's now unlocked by open models.
Will there be an open model that becomes a God-tier model?
I think it's possible.
And we're seeing really exciting developments from ZPU AI and GLM where some of the tasks, they are frontier.
And we all saw with the Kimi model how for web development, it became the best model.
And that sent this new shockwave across the market, which is it's less about a gap and it's more about a head-to-head competition, which I think makes all of this even much more exciting.
This is a bit of a sensitive question, but what do you think about this and the geopolitics around it?
I think, you know, a lot of the geopolitical angles around this start with, you know, where the model's from.
And the more we spend time with customers and users, a lot of it's actually how the model's run, where it's run, how it's run.
Is it run in a secure environment?
And that starts to matter a lot more.
But I do think, you know, look, it's super important that a customer in the US can use a model trained in the US.
And we have two classes of customers we speak to.
One is they don't really care where the model's from, they care about where it's run.
But for every one of those, there's a customer that's saying, I really care about where the model's from because it's data, it's not even just a security issue as much as how does the model speak?
We all go through and communicate and we all go through.
You know, a lot of the models go through phases where they sound more robotic, they sound more friendly, and a lot of that matters too.
But I think the highest order bit is obviously making sure that you have a model that end-to-end, you understand where the data's from, which is great from the Nemotron models.
You can go and introspect what made this model, because if you're putting in a mission critical task, which people are absolutely using open models for mission critical tasks.
There's a post online about how Lama powers the analytics of a power plant to detect surges in Finland to make sure that the lights stay on.
That's where these models, the model origin really matters.
For like critical tasks like that, how do you ensure that a Chinese model?
even if it's hosted in the U.S., isn't basically booby-trapped to cause problems.
The Manchurian candidate problem.
Exactly.
Have there been any known cases of the Manchurian candidate yet?
I think not that I can think of off the top of my head.
I feel like I would have heard about it.
What you don't see a lot on some of the press articles is how robust some of the IT and security teams are.
the businesses that we know of, the top Fortune 500 businesses, they're really used to this already because open source software, if you think the average application, has thousands of dependencies.
This isn't a new problem.
And all it takes is one dependency for there to be a major security issue in the entire application.
Yeah, supply chain poisoning is insane.
It's a thing.
It's been a thing for decades.
And it's not...
new in that sense.
It's a little more opaque because you can't dig into the model.
It is deterministic.
But it's deterministic.
And if you screen the model properly with safety checks, by and large, at least what we're hearing from customers, is that can be solved.
Do you want to talk about the origins of Ollama?
You guys came up through the Docker ecosystem and a lot of people watching would love to be in the position you're in where you have this sort of enduring brand moat that looks like it will extend for really till the end of time.
It's a very powerful situation to be in.
You basically found yourself on top of a giant oil well.
For those out there wildcatting, can you tell us that story?
You were actually working with Jared in 2021.
Yeah, you know, my co-founder and I previously built Docker Desktop while at Docker.
So we really got an understanding of what makes a great developer experience.
But I have to say the first few years of Olam as a company was really in search for what's the right problem to solve with this.
muscle we've built of trying to design a great experience for developers.
You applied to YC with a very different idea, right?
For sure.
Do you remember what the tagline was when you guys applied to YC in Winter 21?
I think it wasn't well defined.
I think we realized, let's go back to building a really great desktop experience for containers and Kubernetes.
I remember what I wrote down on the application.
It was a Kite-matic for Kubernetes or a...
Docker desktop for Kubernetes.
Yeah, which was effectively Docker desktop.
They had a great Kubernetes one.
I think, you know, it's one of the challenges as a second time founder that, you know, Michael and I have told ourselves, we tried to over-engineer the idea in many ways.
And I think even the two to three years after, like Olama, we did YC in 2021 and Olama wasn't launched until July of 2023.
After we raised our Series A, after, obviously after we had done YC, that journey was one of, really in search for a customer problem that could delight a developer.
And in some ways it was almost a good thing that we tried different ideas and pivoted until 2023 because that's when Lama came out and started the open model wave.
Which is why it's called Olama.
Not necessarily.
Okay, no.
Oh, really?
Lama means generally from our experience, whether you think of local Lama, the subreddit, Olama really just stands for open models.
You know, as we were looking through the name, it wasn't necessarily from an existing model.
It's more LLM.
Exactly.
It's like the animal plus LLM.
Yeah.
I think having that character was important.
So we were like, what's a good name for a character, a face you can put to the name?
Because open models are scary.
You need a good animal mascot sometimes.
Docker had one.
GitHub.
He still hasn't taken my advice to have llamas come to actual olama events.
Oh my God.
I'm curious what your Series A pitch was, because you raised from Benchmark, like, fantastic investor, but all of this, the future we're in now hadn't quite taken off in 2023.
So what was, like, the pitch and the vision back then?
Yeah, and we partnered with Benchmark in 2022.
So it was Dolly days, but pre-ChadGBT.
When it came down to the pitch, I think we weighed so much on, like...
hey, we're trying to build this great developer experience.
We're solving this security problem.
And we had known Peter, a partner at Benchmark, from our previous lives building at Docker because he was the Series A investor in Docker.
And so a lot of it was weighted on the people and also why we exist.
I think the what, I mean, solving SSO for Kubernetes, which is a real problem, wasn't really our passion.
I think we were really lucky to find a partner that...
could see us for what we stood for and what we were trying to do versus the point in time, you know, problem we were solving at that point.
I see.
So you raised the A sort of pre-pivot, didn't you?
Correct.
Okay.
I didn't realize that actually.
I was just looking at this cloud tokens by model family graph.
And like, basically, if you just look at this graph, it looks like the Olama story begins in February, 2026.
And it like explodes thereafter, which is like so funny because like, of course, it actually goes back to like 2021.
How was it like to be like sort of lost in the wilderness for like many years working on stuff that was like kind of working, but like not really taking off.
And then all of a sudden to have things like just like explode, like how did, how did it affect you and your co-founder psychology and the team and the employees?
What was the experience like?
It was definitely scary.
And, and, and for a few reasons, you know, one is like when you're, when you're trying to solve a problem for devs or for a customer and you're just.
getting on the phone with them over and over again, and it's not totally clicking, that's, you know, it's less about the, are we in the headlines or is the project taking off the product we're building?
It was just, are we truly actually solving a problem for somebody?
And I think being lost in the wilderness, like what's your North Star, that customers are generally a great North Star, but not seeing the North Star is even scarier, right?
Because often, you know, what problem you want to solve, you just haven't figured out what problem.
And I think the, you know, Michael, my co-friend and I, we started this company because we built a company in the past and we ended up being acquired by Docker very early.
It was just the founding team.
And our North Star was saying, we want to go solve a great experience for developers with something they find really hard.
But man, in the two years where we're just finding that problem, it's really scary.
You know, we had a team of more than 10 people.
which made that really hard.
And I'm so thankful to that team for staying by our side as we went through different ideas, you know, and what's not really obvious is we went from this security for Kubernetes to then like security for developers on the desktop, which is like the pivot that we've never spoken about.
And then we kind of took that form factor when models came out and we said, well, it was a leap, but it was, we knew kind of the kind of problem and the feeling a developer wanted to have, but.
LMs finally made it realize like it was, it was crystal clear at the point when we tried running the Lama model and it was really hard and we're like, okay, this is a problem.
And it's really impressive when you get it working and it's kind of just a zero to one moment.
I'm curious for the story of that pivot, because there are actually like many pivots in the Olama story, but probably like the most critical one was like the pivot to Olama to doing like locally hosted LMs.
How did that come about?
Were you just like tinkering with ideas on the side?
And when you found the idea, was it really obvious to everyone in the company that that was the thing to do?
Or was there like a big debate and it wasn't until it took off that it became clear?
We, you know, sat in a room together.
I remember we were in Toronto because we had a team split across Toronto and Palo Alto and now we're predominantly in Palo Alto.
And we were saying, throw everything out.
Like if we had to start from scratch and we were just joined YC right now, what would we do?
You know, we had seen two big problems because we had talked to some users in LLMs.
We tried using open source LLMs ourselves, which is LLMs in general.
One problem was, could you build a gateway to access any model and host that and make that really seamless?
Back then, we were thinking of it as like the segment for LLMs.
That's a good way to think about it.
Which I think has become really this big router idea, which is only at the beginning.
It's a massive opportunity.
And the other problem was we were a bunch of ex-VMware, ex-Docker folks.
We know how to make things run.
And so let's help make things run with open models.
And then we kind of, yeah, systems.
And so we kind of tried to really introspect our team, which I wish we had done sooner because security is a very different team and sale than developer tools.
And just by doing that, we gravitated towards saying, let's just try this thing.
Let's give ourselves two weeks to launch the first version of Ollama.
And then Llama 2 came out and we said, that was right at the end of the two weeks.
And we said, okay, we're launching it.
And we just had a bias to action.
If you think back, like.
In two weeks, all of that happened, going from idea to shipping it to getting to more users than we had ever had with our previous stuff.
And before that was two years of just, frankly, overthinking the customer, the product, and just not getting something out there.
The first time I actually heard about Olama was on Reddit.
I didn't realize it was you guys.
I was on like that logo.
I just was interested in like running local models.
And it was on like, I think the local LLM.
subreddit or whatever and everyone was just raving about olama and how great it was i was like oh it's a yc company yeah i found out later actually because you you were called a different company you weren't in our internal system yeah i remember catching up with jared and saying oh hey by the way there's yeah you know all that security stuff we have this olama thing now i think you were catching up with jared and then i bumped into you on the stairs and i think you had your t-shirt or some swag or something and i was like like yeah you guys are I was like, meeting a rock star or something.
It would have been really hard to time this, but I wish we had taken that leap much sooner.
I mean, the best time to do it was during YC.
It was impossible.
ILMs didn't exist.
Exactly.
Lama didn't launch yet.
I guess you were also one of the first GitHub projects that very quickly got to 100,000 GitHub stars, right?
Do you remember how long?
It was like very quick.
Yeah, I can't remember exactly how fast, but it was much faster than Docker and Kubernetes.
To your point, things kind of just started working and started taking off.
And you're really, as a founder, just beside yourself because you can't totally explain why.
I think it's the best way to explain product market fit.
And there are different levels of product market fit.
We only started monetizing earlier this year with Ollamas Cloud.
But just to see people fall in love with the product, it's such a zero to one moment that I wish we had done it during YC, but in some ways it wasn't possible.
I also think it's just kind of wild to put into perspective.
Like you sort of went from being in sort of like the cranks on Reddit, like interested in running their own like rigs at home to like 85% of the Fortune 500.
two years or something like that.
That's like a pretty...
That's the homebrew computer club to broad computer adoption, like speed run that took 10 years for the PC.
It took like 18 months, 12 months.
Yeah, and that's one of the things that surprised us the most.
Because I think, look, I think Open Model is the original users, very much hobbyists, just tinkering, oh my God, this is even possible.
But very quickly, because, you know, two things.
One is they were free to get started with.
You could run them anywhere.
That is incredibly helpful to a Fortune 500 IT developer team because they don't have to ask for permission to use it.
And so it just happened what was really good for a hobbyist user translated very quickly to a developer within a business.
It just happened to be a case where that was it.
For example, databases, we saw some of this too where a database that started for devs like MongoDB very quickly also moved to enterprise.
LLMs are stateless.
It made for such an easy transition.
Now, moving to the cloud, there's a lot more in play.
There's an economic question if you're a customer.
There's obviously security.
Where's the model running?
But what's beautiful about open models that both hobbyists and IT developers loved is you could just get started.
You didn't need permission.
Can we talk about the monetization angle?
Because this is interesting too.
2023, Llama 2 takes off.
All of a sudden, you've got all these users, 100,000 GitHub stars.
You've clearly found something.
But it's basically like Reddit cranks who are using it.
You're making no revenue, and there's no obvious path for how you will ever make any revenue from all these cranks on Reddit.
It was two years before you actually figured out a business model for it, which...
funny enough, is exactly the position that Docker was in.
Like, how did you think about it during those two years?
Were you worried about it?
Was the team asking, like, what's the business model going to be?
How did you think about, like, figuring out how to make money from it?
I think there's always two ways that we saw open models being able to monetize in a way that's great for the company, great for the developer, and great for the customer.
And one of them was a privacy-focused AI product, which Olama started really started with that in its open source incarnation.
But we always felt that there was this moment where, you know, you weren't using Lama with the Lama models, for example, with tool calling right away when they came out.
So there were use cases where it was still reserved for the frontier models.
And again, at the risk of overthinking it, we kind of saw that there wasn't the level of product market fit with open models that closed models had.
And in some ways...
philosophically, we want to align with when that happens, we want to be there to capture that.
I think it happened this year with coding agents running with open models because you had the largest consumption of AI being matched with finally open models being able to service that.
There are a lot of opportunities along the way to do it privately, securely.
Again, a lot of the Fortune 500 have already adopted Lama.
But we really asked ourselves, what would be the most important problem we could solve for a customer?
And the local piece, while an important part of that story, never felt like the whole story, which was, How do you access open models for the hardest problems?
And so in some ways waiting, we knew we had to wait a little bit for the market to mature.
At the same time, what are the risks of waiting?
Well, you build a culture, if not careful, and we had learned a lot of this from our Docker days, where you don't think about monetization.
It's not a priority.
I think from our previous battle scars as a team, we kind of had, we knew about that.
But I think the other component, which is really important is making sure you keep in touch with your customers.
Like one of the biggest risks of having an open source project that takes off is you consider your user base and customer base, like your customer, just a blob on the internet, which is a really risky way to think about customers because you want to meet them, figure out their needs.
What are they doing?
What do they want to do in six months?
What's their story?
And I think that's the thing I wish we had done a little more in the last few years.
And we're doing a ton of that now.
One thing I'm curious, when you went through YC, you guys were second time founders.
I'm curious what got you to decide to do YC, actually.
You know, we went back and forth on this for a lot, which we shouldn't have.
We should have just said, of course, we're doing YC.
But by and large, starting a company is a really lonely experience, even if you have a great co-founder.
And Michael, my co-founder, was the co-founder of my first company.
He was my college roommate at University of Waterloo.
But it's still lonely.
And I think just having a set of peers, even though we did it during the pandemic, just talking to Jared and like...
five other groups of founders every week really helped you feel less lonely.
And I think that's such an important part of it.
And then of course, when we finally moved down here and there was no more COVID, the network was just incredible.
And the fact that we could meet founders building on open models, building on any kind of AI, you know, we kind of knew that was going to happen because we had known so many founders from the University of Waterloo who had done YC pre-COVID.
And they were like, it's really about getting together.
And like, that was a big part of it.
And we knew that was there.
And, you know, I think that's what made it a no brainer.
But also just, I think there are a lot of mistakes you can repeat that you don't have to.
And what I love about the YC community is how transparent founders are with each other about those.
And, you know, I still keep in touch with the founder of Docker, who's an investor in our company.
And we're able to talk about some of these challenges we saw in the previous generation of companies that, you know, don't necessarily have to repeat or things that worked and we can bring into the future.
Yeah, if you just don't repeat one of those mistakes, that sometimes is the mistake that would have killed the company.
Potentially, yeah.
Our famous saying, a bunch of our team is from companies that ended up working great and Docker's doing phenomenal now, but whether it's some of our team was early, early at VMware.
And there are always ups and downs.
And I think...
just having a group of people around the table who have a collection of those and also what worked.
Actually, what worked is actually even more important.
And just being able to like have that muscle memory is a big part of it.
Oh man, I was just thinking about this because we obviously hang out with...
work with a lot of 18 year olds or 19 year olds.
And then sometimes they're always asking like, well, what should I do?
And then I'm starting to realize like one of the more important things is if you've never worked on a team that shipped really amazing technology to like a lot of people or like, like just real clear product market fit, like do that once, like even if it's a month, even if it's like three months, you would learn more in those three months because then you know what good looks like.
And then without that, it's like, I mean, it's not like it's impossible.
Like people.
at YC do figure it out because, you know, but it's that much harder.
Like the difference between having seen something that actually works from like beginning to like some form of like, this is what the bug database looks like.
And this is how we release.
And this is the quality that's necessary.
And here's like the bar that we hold each other to.
Having seen that, it just like multiplies the chance that people succeed.
So it makes sense that.
you know, starting off with a co-founding team that has seen a lot of that pretty powerful.
Yeah, I think it provides you a set of values you can work around, especially when you have so much power in your hands with AI.
There are just parts of it that AI can help you with, but it won't hold you accountable to it.
And how does software work?
And to look at Ollama, for example, I'm sure there's versions of Ollama running in the wild from two years ago.
How will your software work when somebody...
falls in love with it and continues using it for two years, is it still going to be working well?
Hopefully they update to the latest software or it's a cloud service.
But I think you build that muscle memory.
And we definitely have that from a lot of our more senior engineers on the team who are at VMware or Nysera, for example.
But at the same time, I think there are a lot of lessons we learned in the previous generation of DevOps and infrastructure that aren't valid anymore in the AI world.
Oh yeah, tell us about it.
What have you found?
What is not valid anymore?
I think a good example.
that I classically used is there was this generation of companies called platform as a service.
And the Heroku's of the world was a great example of this.
I mean, Docker started out as that.
Docker started out as a platform as a service.
And there's this concept that if you're a layer on top of something else, that you're in kind of a vulnerable position as a startup, which is absolutely not true in the iWorld.
And in fact, going up the stack.
can sometimes be even better because you're closer to the customer.
In an infrastructure world, that's also the case.
And that was like an analogy that we had to like, so many of these muscles, we actually had to break building Olama.
Another one was, you know, these LMs are never perfect.
And like in the systems world, you want everything to be exactly as it's designed to run.
It's tested, it's validated, but LMs by definition are not.
That's a feature, not a bug.
Exactly.
It's a feature.
You want it to be a little non-deterministic, I suppose.
And I think from building a team too, it's.
that you know with ai now they're just problems that you don't need to staff as heavily whereas you you you did 10 years ago right if you think about what does your customer support pipeline look like um what does it look like to deliver a cloud service like it's a very different world with ai because how do you build a service where no engineer knows exactly how all the code works which is obviously the case now and so there's just new lessons we're learning going from like a you know, some of our team from infrastructure 1.0 in the 2000s to cloud in the 2010s to now the AI space, there are a lot of rules that break.
I mean, you're probably actually doing an incredible service to like both sides of the ecosystem and that like the end users get this like very clean.
thing that just works, especially like the tokens just come out and they're very clean and the API makes sense and it's rational and logical.
And then on the flip side, like, I mean, if you don't have a layer like Olama, I've directly experienced this where it's like, oh yeah, the underlying inference provider has a weird error for, you know, if you put this parameter in this way or it expects JSON and, you know, it's not documented.
It's just like this insane minefield.
Like, you know, the agents can kind of figure it out, but like, you're going to like bang your head into the wall for like a couple of years.
couple hours before, you know, the agent figures it out.
And in the meantime, you're like, this is a terrible experience, you know?
And so you're like in there probably helping the inference providers fix all these fundamental bugs too.
Yeah.
And it's part of the job we do.
And I think one of the big opportunities in the open model landscape is curation and taking a fragmented universe of models and inference technology and cloud services and harnesses and like making that actually.
just work is a really valuable problem because the end developer, to your point, they just want to build their software, right?
They just want to build stuff.
They want to build their next company, their next application.
And I think that's where we come in, but it's where a ton of great services also come in.
And we saw, you know, open router obviously is a good example of that from a wide model selection.
So the developer doesn't have to sign up for, you know, a hundred different.
providers, they can just go to one, they can pay in one place.
I think we've seen with open code, you know, you can have one harness that integrates with any model.
It's a really powerful experience for developers just looking to try the next model to see if it solves their use case better.
So this curation and, you know, when there's a abundance of models and providers, now there's a scarcity in bringing that together into something that works.
Thank you so much for joining us.
That's all we have time for.
Thank you guys for having me.
