# Open AI Models Reshape Pricing, Infrastructure, and Market Moats

**Podcast:** a16z Podcast
**Published:** 2026-07-24

## Transcript

If you kind of bring it back to very business-first principles, if you're providing a product of value, capitalism will find a way to make the supply chain work for you.
So if you have an open-made model that is providing value, that means that every part of the stack underneath, whether it is a neocloud, the chip provider, somebody who provides gas turbines or fire suppression, is going to orient itself to provide value.
If you're providing a product of value, capitalism will take care of all the rest.
If you go look at how the rest of the ecosystem is doing, the growth is pretty strong and spectacular.
And I think you're going to see that continue.
Open source AI is moving faster than ever.
And the balance of power in the industry may be shifting.
In this episode, Theo Jaffe and Sophia Puccini are joined by former White House AI policy advisor Sriram Krishnan.
To unpack with the latest wave of open models means for frontier labs, AI policy, pricing, cybersecurity, and America's position in the global AI race.
We are back.
We are live with Shriram Krishnan, who just finished his tenure as the Senior White House Policy Advisor on Artificial Intelligence.
Previously, he was a general partner at Andreessen Horowitz and held senior roles at Microsoft, Meta, Snap, and Twitter.
So, Shriram, we're so glad to have you on.
Welcome to MTS.
Thank you.
I've been a fan of everything you folks have been doing for the last few months and excited to be here.
I think this is the first time in about two years I've been able to do a video appearance without a suit and tie on.
So I am so excited to be out of that.
Yeah, yeah.
So there's so much going on in open source last week.
We just discussed we had...
GrokBuild was open source, and then we had Thinking Machines, and then KimiK3, and then Quen 3.8.
So you tweeted the other day, KimiK3 is a big moment with multiple implications for the entire industry.
Could you go into a little more detail on that?
What are these implications?
Yeah, so if you go back maybe four or five months, I think there was a moment in time when the only leading models were, I think, Opus 4.
six or four seven at the time gpt54 or 55 or wherever we were and it felt like there was really no one else and we were on this curve or some self-improvement where the frontier labs were really going to draw really far away from everyone else uh i think the last few weeks uh if you are in the token consumption business uh which i am and i think many of uh you and your viewers are it's been a great time because let's see You had Elon and Michael at Cursor, you know, the SpaceX XAI team come out with Glock 45, which I've been using.
It's a fantastic model.
I think we forgot to mention this.
We had Alex Wang and Meta come out with Muse Spark.
which is also awesome.
Last week, we had Mira Thinki come out with Inkling.
I don't know their version number, but the first version of that model, which I think is nearly SOTA on many, many benchmarks.
But I think that the big news of the last three, four days was obviously Kimi K3 coming out, I think on Thursday or Friday.
And then I think the last 24 hours.
I haven't really played with it yet, but Quinn coming out.
So there's a lot of choices and alternatives coming out.
And what I was referring to is, with Kimi K3, it's the following.
One is that it's just great to have choice in the ecosystem and to be able to point your harness of choice or your agent of choice to multiple models.
Second, I think we are in this really weird moment now where some of the American frontier models are constrained.
For example, on cyber and on security.
And I was talking to a friend of mine where this person was actually starting to do security work.
using Kimi K3 rather than Fable because with Fable, he would run into these refusals and safeguards.
So that seems like a very weird spot to be, which we can talk about for a second.
I think it's also, you know, it's probably inevitable that if you are having choices from where you get your intelligence tokens from.
That's going to put pricing pressure on the frontier models, which means I think you'll probably see the frontier labs have to drop token prices or find ways to match pricing, which means it's probably going to erode into their gross margin.
It's probably great for the NeoClouds and every other layer of the stack because if you're a NeoCloud like...
a base 10 or a fireworks, or if you just have a bunch of black walls and you can power it in, you can run that and, you know, you can capture some of the economics, which is probably going to go to the frontier lab.
So I think it's great as a consumer.
I think it's great for the ecosystem.
Lots of obviously questions on security, on distillation, on, you know, how this is good.
But a very, very interesting moment.
Totally.
So I guess like the first question here would be, where do you predict?
the frontier labs, like how do you predict the frontier labs are going to react to this?
Like now, you know, the ecosystem has been sort of changed.
Like there's no going back.
Like Kimmy has been released and it's very close to the capabilities of Fable.
So what do you think is next for like an Anthropic or an OpenAI?
So a few things.
I think it's very clear the frontier labs are going to push at the very, very frontier of, you know, the jagged performance we get.
You know, as somebody who spent close to the last 18, 19 months trying to make sure America wins.
I really want to see the American models, whether it's closed or open weight at the frontier, which, and so I think they'll continue to be that.
What I suspect these open models could really start putting pressure on them is one on pricing, because it may turn out that the number of tasks that you need absolutely frontier intelligence from is Let's call it one subset.
But for a lot of other tasks, for example, I have an agent which checks my email or I have an agent which quickly scans through my calendar.
You may not need Frontier tokens.
You may be able to get by with Frontier minus one or your OpenWay.
token of choice.
In which case, I think you're going to start seeing pricing pressure come on the frontier labs.
You're probably going to see them try and make sure the absolute frontier models stay available and accessible.
I think we've already seen Anthropic, and I have no inside confirmation about this as a motivation, but you've already seen them extend Fable.
I think Fable is originally supposed to be available until...
I don't know, like a week ago, it's already been extended.
I predict there'll probably be more extensions just because otherwise you have an open model, which is very much near the frontier.
I think the other interesting question is, where is the real moat?
if you're a frontier lab, is it in the intelligence or is it in the harness?
And I think, you know, Cloud, you know, Cowork and, you know, Cloud Code and Codex are amazing products.
And I suspect you'll probably see like more effort on making them sticky because it might be that, you know, the actual intelligence itself or at least the...
the levels behind the actual frontier are much more of a commodity.
So what does this mean for the ecosystem?
One, this could put some revenue pressure on the frontier labs because I suspect that people might start pointing their harnesses at some of these other open-weight models for any number of reasons which we can get into.
I think it's probably great, like I said, for the near clouds.
and anybody else with power and GPUs.
And I know it's going to be a very, very interesting time.
Right.
So there's been, speaking of open source, there's been some chatter.
Axios just reported this morning that the Trump administration is considering restricting Chinese open source models, either because they are Chinese and might be a national security issue, or because they have reached like Fable level capabilities and agentic coding, and it could be a cybersecurity issue.
So do you think the U.S.
government will do things that will crack down on open models?
And do you think that there's anything that they should be doing in this respect?
Well, I have no inside information.
As I like to tell people, I am no longer living off your American taxpayer money.
But I have a lot of friends there.
And I think...
From everything I hear, if you go back to a year and a half ago, in the first few weeks, the Trump administration said, hey, we're going to come out with the AI action plan.
And when we came out with the AI action plan, it's right in the first page.
If you're going to go to the first section of it, you're going to see the action plan talk about how important open source is.
Now, I think the thing to think about is there are multiple different issues.
Number one, I don't think it is great.
that the leading open-weight models or open-source models are not American.
I think there's some great innovation happening with Moonshot and with DeepSeek and with Quen, but I would much rather prefer that the leading models are American.
And by the way, there's lots of great efforts happening.
I expect them to improve.
You have Gemma from...
Google, you have Nemotron from NVIDIA, you have obviously Thinky, you have newer startups like Reflection coming out.
I think they will continue to be better.
But we are in a moment of time when the leading models are Chinese, which I think is not great.
I would much rather prefer them to be American.
I think the second part of it is I grew up loving open source.
Open source is a big part of how I got into computers, a big part of my career.
And I was a big fan of Linus's law, as in Linus Torvalds of Linux fame's law.
And his law was that given enough eyes, all bugs are shallow.
And what I believe with that is that open-weight models are inherently secure because, you know, when you download a model of hugging face, it means you have the entire world being able to take it apart, inspect it.
you know, fine tune it, modify it, look at it in ways that you absolutely cannot if they are closed.
So I kind of believe that they bring a very, very different, positive angle to security.
The weird, I think, or the moment we are in, which I think is not great, is I think it was in...
incident with Hugging Face that got reported on earlier today, which, you know, I think it was an active tweet.
And what Hugging Face seems to have found is that somebody was using, you know, essentially definitely an AI LLM agent to hammer them at multiple places and to kind of like, you know, find ways to break in.
And, you know, I think the way to counter that is to make sure American or Western or allies defenders have access to the best models to make one, our software, like just more secure.
So which is why, you know, like I said, I don't love that we are in a moment where it may be harder to look at your own code for exploits with, say, Fable than it is when you use a Chinese model.
That's just like a weird spot to be.
And I think, you know, we should fix that.
Now, one very interesting discussion, which has been happening a lot on X, is about distillation and what that means.
So it's kind of a complex topic because there's a few things in there.
So first of all, every model we have today distilled off of all human knowledge, right?
Like if you go back to the original GPT or...
or the original cloud, you know, they all had to derive of crawling off the internet, you know, crawling off all of our, you know, blogs or tweets and content.
And, you know, they kind of consume human knowledge and to bootstrap that.
So distillation has always been a core part of how these models have been trained.
Second, if you, you know, okay, let me ask you this.
When I write a tweet these days, I am terrified of accidentally using multiple hyphens or accidentally saying something which will cause Pangram to say this is AI generated.
I once wrote a tweet recently, which I put it through Pangram to make sure it felt human, even though I knew I had written every human generated token.
And so what we've all seen is, I would say, the rise of AI slop on the internet.
So if you think about that, The internet has grown a lot over the last couple of years.
A lot of that has been AI generated.
You and I see content every single day.
We're like, well, there's definitely some tokens in there.
And that also goes into the training of these models.
So distillation of AI models is a core part of how these models are trained today.
There's just no escaping that.
That's number one.
Number two, I think the issue I think people are talking about is that it, you know, Some of these models may be distilling at scale in kind of these industrialized ways where you have maybe a bunch of fake accounts, or maybe you have people reselling, say, cloud subscriptions to somebody else and routing it, any number of ways which are breaking terms of service and which are bad.
So here's what I think.
Number one is I think all the labs, you know, with the help of, you know, everybody, you know, from the government should be doing everything they can to make sure, you know, whether they're doing KYC, whether...
They are doing things to check the IP addresses where they're coming from.
There are a lot more sophisticated things they do.
They are making sure this heavy industrial extraction of reasoning traces isn't happening.
And I'm going to steal this from Dean Mayer of Sequoia, who had a fantastic post.
And please give him credit for this in his tweet.
He wrote this yesterday.
I think the situation which is bad today...
is that some of these models from other countries can train off American models.
Whereas if you are an American open-made model, it may be really confusing or challenging on whether you can distill off of other American models, right?
So if you kind of have a really uneven ecosystem here, where if you're a Chinese model, you could probably get a bunch of reasoning traces.
But if you are, say, a new Valley startup, and you want to use some reasoning traces, you don't know what the legal situation is.
So one great idea, which I think came from him, I think...
Ben Thompson of Stratachiri had a similar idea today, was to basically say, how do we find a way to make distillation acceptable in any number of ways?
Whether you are getting outputs of other models, or sometimes it's more subtle.
If you look at any American open source model today, they are using Chinese models as a teacher, or in a way as a part of the fine-tuning process.
And how do we make sure that is protected and enshrined?
So if you're an American...
in a model company, you have the same level playing field as the Chinese models.
Now, again, I think we're in this kind of this very interesting, unique moment in time.
But overall, I think these open-weight models are great for the ecosystem.
It's more choice.
It's people competing on price.
It's all the great things that innovation is supposed to bring.
Right, totally.
So another consideration that is going on right now is the idea of AI being able to automate to improve first and then automate AI research.
So, you know, Anthropic has written about this.
OpenAI has talked about, you know, building in first an automated research intern and then working towards automating AI research.
And many people believe that if that were to happen, the pace of AI capabilities progress would go very, very, very fast.
So what kind of policies do you think the government, if any, should have around this?
What do you think the government?
is going to do around a potential very rapid increase in capabilities of AI, if that were to happen in the future.
Well, again, I don't speak for the government anymore.
You know, I don't have any inside information.
I don't know whether the government has a real role in this.
Like, look, whether RSI is real or not, where you are on the exponent is a much debated topic.
You know, I've heard many, many schools of thought where, you know, they believe you're going to have automated AI researchers in a couple of years.
But then there are others who think, well, there are fundamental improvements that you cannot crack.
And yes, you're going to see improvements from models being able to train other models.
But the exponent is going to be a lot more gradual rather than a very, very, you know, steep.
curve.
So I think when I was in government, for me or for David, the entire focus was how do we have a fantastic ecosystem where there are people competing to build products, you know, you're pushing innovation as fast as possible.
And then when there are credible risks, you know, making sure when they appear, making sure you tackle them.
So for example, with cyber, when you have a credible threat, when you know these models are capable of, for example, generating exploits in the latest firmware or the latest operating system, how do you then go tackle it?
So when somebody gives me this question, I sometimes find it very theoretical, and I often find it a lot more useful to break it down into, okay, one, how do we make sure we have a bunch of competing choices?
Two is, what do we know to be credible?
And, you know, how do we actually particularly tackle that?
I think there are very credible threats on cyber, on biological advances, a couple of other topics.
And I think...
There are definitely efforts to try and tackle just those.
And then, you know, and I think the final part I would say is that in a lot of ways, I'm a big believer in using AI to help.
So often, for example, the answer to helping with cyber is if you have, you know, using more AI to go scan your code base and making your code base more secure.
So very often, I think, like, how do we then deploy to actually meet these threats as we go along?
Totally, totally.
So another point that Dean Ball raised actually was sort of this sentiment that like open weight models deter CapEx and they kind of like destroy the ability for frontier labs to monetize their capability leads.
So how do we like promote a very healthy like open source ecosystem while, you know, ensuring that the frontier labs are still able to like do what they're best at?
Well, I think at the end of the day, If you kind of bring it back to very business first principles, if you're providing a product of value, capitalism will find a way to make the supply chain work for you.
So if you have an open rate model that is providing value, that means that every part of the stack underneath, whether it is a new cloud.
whether it is a data center, whether it is a chip provider, whether it is somebody who provides gas turbines or fire suppression to a data center, every other part of the stack is going to orient itself to provide value because you, as an open-weight model provider, you might be working with a bank because they don't want to work with a frontier model.
They want to work with somebody that they can run in-house.
Or if you go look at what somebody like...
think he is doing, they are fine-tuning models for specific clients using internal data that they don't expose.
So, you know, so I kind of think of like, you know, if you're providing a product of value, the capitalism will take care of all the rest.
And by the way, I think I'm being borne out.
Like if you go look at the fraction of...
every other part of the stack.
If you go look at how the inference clouds are doing, if you go look at how the rest of the ecosystem is going, the growth is pretty strong and spectacular.
And I think you're going to continue seeing that continue.
So to the extent that you can tell us, what are you working on now that you're out of government?
Well, I'm going to be doing more, you know, live drop-ins, I think.
Streamer.
No, look, I think...
Well, look, I had this amazing life experience, which is unparalleled, and it was such a unique honor.
And it's kind of really, you know, given me a sense of what countries and companies need to do to work with each other.
and to make sure everybody gets intelligence, whether it's from the frontier labs and open-weight models or however that happens.
And I want to try and make that happen in some shape or form.
So I'm going to be annoyingly elusive, but I think that mission of making sure America, our allies, get access to AI at scale.
with having governments and these companies work together is a very important one.
I've been working on it in many different ways, if you think about it, for many years now.
So I want to continue to do just that.
And write some banger posts and jump in on live streams from time to time.
Yeah, well, we're excited.
This has been so great.
Thank you so much, Shreem, for coming on MTS.
Yes, I'm sure we'll have you on again.
Thank you.
Thanks for having me.
Much more to come.
Absolutely.
Thanks again for listening, and I'll see you in the next episode.
This information is for educational purposes only and is not a recommendation to buy, hold, or sell any investment or financial product.
This podcast has been produced by a third party and may include paid promotional advertisements, other company references, and individuals unaffiliated with A16Z.
Such advertisements, companies, and individuals are not endorsed by AH Capital Management LLC, A16Z, or any of its affiliates.
Information is from sources deemed reliable on the date of publication, but A16Z does not guarantee its accuracy.
