# AI Infrastructure Economics and Strategic Shifts

**Podcast:** Last Week in AI
**Published:** 2026-08-31

## Transcript

Hello and welcome to the Last Week in AI podcast where you can hear a chat about what's going on with AI.
As usual in this episode, you will summarize and discuss some of last week's most interesting AI news.
As somewhat usual, we'll be also including a bit of news beyond last week since we are catching up.
I am one of your regular hosts, Andrei Kurenkov.
I studied AI in grad school and now work at the startup Astrocade.
And I'm your other regular co-host, Jeremy.
Yeah, Gladstone AI, AI National Security stuff.
Soon to be doing a series of like more public work around investigating kind of the AI end game and national security and lab and infrastructure side of it.
But yeah, and I apologize.
The absence last week was on me.
We had our big sort of family green card activation trip.
And so anyway, traveling with a baby to LA for like some work stuff and some pleasure stuff.
So there was a lot going on.
Things kind of got hairy there.
So I appreciate everyone's patience.
We should be more settled.
I was just telling Andre, like I've got one more trip coming up the week of September 17th.
So there may be a hiccup there.
Andre may be able to get a co-host for that.
But just to kind of put that on everybody's radar, things should be settling in more now that all the green card insanity, and it is insanity, is over.
Yeah, looking forward to the next phase of the story.
Yeah, it's a bit of a process with these things.
And as you said, travel can be tricky.
Luckily, we didn't miss too much.
Missing last week relative to before, things kind of settled down a little bit since our last episode.
So nothing too crazy to cover as a quick preview of this episode.
We'll be catching up a little bit.
There have been some new models and like tool updates to mention, but nothing too big.
We will be discussing things like Gemini 3.7 and GROC 4.6, some new model releases, but nothing sort of earth shattering.
Applications in business, again, nothing like huge, some interesting developments and some slightly more kind of nuanced things that if you're following the industry, it'll be interesting.
Policy and safety, once again, will be the big section and we'll have a real mix of stuff there we'll be talking about.
drones, which we haven't touched on in quite a while, the security updates from OpenAI and their kind of proposed or supposed slowdown of model developments, which I think is probably the most interesting development.
And beyond that, just, yeah, a real variety of stuff.
And we will try to get to some research as well by then.
So it should be a pretty good mixed episode.
This episode is brought to you by OutShift, Cisco's incubation engine.
Today's AI engines operate in silos, limiting their true potential.
We focus on building bigger, smarter models, but scaling up is just one approach.
To reach superintelligence together, we need to do more.
We need to scale out.
And we actually have a blueprint from 70,000 years ago.
Humans didn't just get smarter individually.
The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation.
Agents are now at the same inflection point.
They can connect, but they can't think together.
That's why OutShift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence.
By creating an open, interoperable infrastructure, OutShift is enabling agents and humans to share intent, context, and reasoning.
The cognitive evolution for agents is here.
Explore Internet of Cognition at OutShift.com.
That's OutShift.com.
We'd like to thank Box for being a sponsor.
The key to unlocking the power of AI isn't in an LLM or an agent, it's in the content stored in files across your company.
Data from Box's annual State of AI in the Enterprise report in 2006 say that 96% of organizations say agents need access to company-specific content, but only 36% have connected agents to trusted content across many use cases.
And that's where Box comes in as a secure bridge between enterprise content and AI.
Again, according to the report, only 36% of organizations have formal standards governing how agents access company data.
76% of respondents say their current AI governance requirements are slowing their ability to deploy agentic AI, and 93% agree that better governance would help them move faster over time.
Box provides agent guardrails, prompt injection detection capabilities, MCP guardrails, classification-based access policies, agent activity oversight capabilities and more to help you connect your ai to your data in a secure way so if you think seriously about adapting ai think beyond the model your business lives in your content and box helps you bring that content securely into the ai era learn more at box.com ai or in the link in the description Before we kick it off, do want to acknowledge we had some lovely comments on the latest episode.
Just picking one out from Golden Dawn had some nice comments and asked that we would discuss a little bit about QEN 3.828B, which is a fairly small parameter model, but did well on benchmarks, like surprisingly well.
And also discuss the comeback of SpaceX AI.
So I did take a look at that and we do have those stories coming up.
So keep tuned to see our discussions.
And so going to tools and apps.
First up, we've got Google announces Gemini 3.7 Flash just three weeks after previous release.
So this is what they are.
terming the new workhorse model, replacing 3.6 Flash.
As it said, these kind of came together very close apart.
A bit of a weird move in some ways because U-Mine has been relatively slow in releasing models this year relative to OpenAI and Anthropic.
We've gone through the whole GP5 family of 5.1, 5.2.
It feels like pretty quickly Anthropic has been through Claude.
4.5 to 5 in a relatively short time span.
But with this latest release, pretty strong benchmark numbers for this kind of quicker, smaller model, not comparable to things like Opus or Sol from OpenAI, of course.
So this, I think, is showing that Gemini and DeepMind are continuing to focus more on the models they need for product.
for the things that are integrated into drive and various products, right?
If you have LLMs doing stuff everywhere, you're not going to be doing that with Gemini Pro, you're going to be doing that with Flash.
And so there's clearly, it feels like a lot of engineering manpower directed towards optimizing Flash and making it cost efficient and quick and capable in a very kind of...
practical move that does to some extent make it so Google is not competing on the frontier of agentic coding and sort of intelligence, at least seemingly in recent months.
Yeah.
And I think that's actually in some ways the headline.
They've shown they can ship fast, but very conspicuously, the thing that they're shipping fast is not an actual frontier model.
And that tells us a lot, right?
I mean, we've been hearing about Gemini 3.5 Pro.
It's coming soon.
That's what they said in May, in July.
Now we're in August, still coming at some point.
So I think it's starting to raise some real questions and doubts about whether Google has the chops to be a frontier lab.
There's a lot about Google's big picture strategy that suggests that they're leaning harder and harder into becoming a, I was going to say cloud, a Neo cloud, somewhere in between, basically shipping TPUs instead of GPUs to compete with NVIDIA, which in the long run, I mean, I think is a pretty risky move.
You know, you risk.
Locking in is the infrastructure layer without the kind of visibility into the needs of the model development layer that you would have if you were a frontier lab.
I mean, we're already seeing OpenAI.
We might talk about that later today, if not in the next episode.
But AI has come out with their own custom silicon, and it's good, right?
Like the jalapeno seems like a legit chip.
Anthropic seems like they have the chops to do the same thing.
But what it starts to look like, and we were talking about this years ago, the idea that...
having a frontier model company that can do hardware may actually turn out to be easier than having a hardware company that does frontier modeling.
And I think that was pretty contrarian at the time.
Now the argument for that seems to be getting stronger and stronger.
NVIDIA's positioning in the market is not exactly weak, but it definitely is going to see viable competition at the design level from OpenAI to minimum, it seems, if the benchmarks hold up, possibly from Anthropic as well.
Yes, Google has TPUs.
Those are amazing.
But as you start to sacrifice frontier model, like genuine frontier model in-house capability by losing Jeff Dean, by losing Demis, by losing all these key players, eventually that stuff is going to erode.
And you'll see something that like similar to what happened to IBM, where you get sucked into these short term rewards where it's like, yes, you can make a ton of money today by focusing on the infrastructure layer.
But at the cost of losing your focus on long-term strategy, where are the insights coming from?
It's from the frontier modeling side that back propagates to the sort of hardware engineering side.
And it's both, right?
They co-evolve and you co-optimize, but you can't do it by flying blind, by outsourcing your partnerships to external labs, which is a big part of the reason why Amazon is not top shelf.
It's a big part of the reason why SpaceX AI had to acquire Cursor.
Like all the stuff points to you need the modeling in-house as much as possible to be competitive in long-term.
That being said, this is a really good model for its tier.
There's no question, as you said, I mean, so the comparison here is more with like a sonnet class or to use the opening eye classification tiers like a Terra.
They have Sol, Terra, and Luna, right?
So Terra kind of that middle ground, again, sonnet between Opus and Haiku.
So this is like a mid-tier model and is very good, it seems, at least based on the benchmarks on that basis.
It is just like a refinement of 3.6 flash rather than a new base model.
I think you might've mentioned that anyway.
It's based on feedback from customers, supposedly in like more RL.
So you look at these benchmarks, some of them did jump a lot, like DeepSwee, so a software engineering benchmark went up from 49 to 65%.
That's a pretty big jump.
Automation bench for multi-step tasks up from 17 to 30%, which is actually quite good.
Still fails in seven out of 10 multi-step automation tasks.
So, you know, that's just the class of model that it's in.
But nonetheless.
compared to models in its class, it holds up well.
The real question, of course, is like, is this kind of new focus on the sonnet tier, the terra tier, and not the sol tier and the opus tier?
Is this indicative of kind of an admission that we're sort of, if not throwing in the towel, at least de-emphasizing the true frontier of capabilities?
Right.
And typically when you say frontier, we mean basically the most capable models, most intelligent models, however you want to define it.
You could make an argument that it is pushing the frontier of the Pareto frontier.
The Pareto frontier, yes.
Yes.
So if you look at some ways of measuring things like the artificial analysis intelligent index versus time per task, and you graph that, you can make the argument that this is actually a leading model on that measure.
It's actually quite fast.
That's part of what they're doing here.
Clearly, it's optimizing not just for small model capability and cost, but also for speed.
So relative to even GBP 5.6 Luna, it might be faster for many tasks.
And again, I think this is driven very much by pragmatic product needs.
And it is a real, I think for me, it's less about a capability question of whether DeepMind can compete or Google can compete head to head on kind of agentic coding and business needs versus what appears to be their current strategy, which is continuing to focus on the model needs for their kind of whole suite of products.
And in particular, things like Google Drive, like you have spreadsheets or you have emails or you even have a presumably search where these kinds of things come in with AI mode.
So as usual, I think with Google as a giant bureaucracy, it feels like this is a corporate level kind of prioritization question of there's more resources being shuffled towards what the...
product requires and less towards this kind of next generation frontier of intelligence kind of effort.
Whether that's a smart strategy or not, you can argue either way.
I will say, I think, if nothing else, having one focus over trying to do everything at the same time is not too bad a strategy.
And they do seem to be doing well on the side of having a good model that can be plugged in to Google Drive.
you know, AI mode and elsewhere.
So Gemini 3.7 Flash for the category, Google is continuing to lead there.
Yeah, I'm pretty skeptical about this approach personally.
Could turn out to be wrong, of course, but I do think like we live in a world where it seems like margins for the frontier models have been going up and value capture at that part of the chain has been increasing quite significantly.
And it seems like there's a secular trend in that direction for a whole bunch of reasons that...
We don't really, we should do a deep dive on the economics of this stuff at some point.
But I think there's a world where this is their IBM moment.
And they're essentially kind of the same way that IBM basically became a glorified consulting company.
And like, yes, you can make a lot more money in the short term by doing that, but you sacrifice doing the hard thing.
And the hard thing is what sets you up for the future.
Even when things look down, you know, everybody was talking about models hitting a wall, you know, a couple months ago.
You know, the companies that kept powering through and kept prioritizing frontier models, of course, they were seeing internally that the whole scaling is hitting a wall story was complete bunk, which we also talked about at the time.
But nonetheless, I'm pretty concerned for Google on this one.
I think it is.
They're just like such a huge pot of money in the short term chasing infrastructure.
And in the long term, so much margin potentially coming in on the modeling side, so much advanced warning and so much flexibility.
If you overbuild your.
I'll pause there because we got to get on.
We can talk a lot about strategy and implications and so on, but that's probably enough.
And to be fair, we've been here before.
Like, who knows?
DeepMind could just release Gemini 4 Pro and it'll be crazy.
Like 2025, they had a surprising comeback story.
So we'll see.
They do have still a lot of talent.
So let's not rule them out yet.
Next up.
We've got Space XAI releases Grok 4.6, a 500k context frontier model tuned for long-running agents coding and knowledge work.
So this is a bit of a catch-up.
This happened, I think, a little while ago prior to last week.
This is an incremental development.
So it's the same base model as Grok 4.5 rather than a larger base model and it continues.
The seeming development of ever since XAI, now SpaceX AI, acquired Cursor, they managed to release new Grok models that are good for coding.
For quite a while, Grok ceased to release anything.
Maybe because the entire exec team left and the company kind of maybe got rebirthed.
It's hard to say, maybe.
And then XAI became a NeoCloud, just giving out computes.
So what happened?
Now we are seeing a GROC 4.5 and a GROC 4.6.
On the benchmarks, it isn't necessarily for Frontier, but it looks to be pretty competitive with...
a GPU 5.6 and with Claude 5.
Being very cost efficient as well.
So that is an honorable kind of trade-off.
I forget the exact numbers, but it maybe is around half the cost.
It's like a pretty competitive model on price.
I will say the 500k context of it is worth noting.
That is when you do Agenda coding.
having half the context is a big limitation.
So one of the issues with the primary benchmarks we see for coding is the ones that are primarily being showcased, like TerminalBench is one and DeepSbyUE.
We haven't kind of gone to a point where there is a leading benchmark for long-running agenda coding.
I would assume partially because it's just a very hard thing to do.
You need a very large-scale project, et cetera, et cetera.
So whenever we talk about new model releases and benchmark scores, my mental model is the software benchmarking is not at the point of being able to really stress test capabilities of the sorts of things you'd be using Cloud Code or Codex for, which is long-running five-hour work tasks with hundreds of thousands or tens of thousands of lines of code in repositories, etc.
Like you have these smaller things for the most part of like fixing a bug in a workspace and maybe a big repo, but it's still a pretty localized change, not developing a new feature which requires some design and testing and so on and so on.
To summarize, Grok 4.6 seems pretty capable and pretty cost competitive.
If people were considering Grok, it could perhaps start to be...
competitive with Cloud Code and Codex.
As far as I've seen, nobody is really thinking, oh, maybe we should switch to Grok build from Cloud Code and Codex.
I'm not sure that OpenAI and Frodovic are too worried about Grok for now.
Yeah, I think the for now thing is maybe the key caveat here.
We've got Grok to go from basically irrelevant to...
Now, yeah, is it a frontier?
It's in the cluster of frontier models, right?
And so you'll be able to find the odd benchmark where it does surprisingly well.
And maybe it is actually the best model for that use case.
It is, as you said, partly meant to be highly cost competitive.
And so that's kind of one of its sharper areas.
The big picture here really is the proof of thesis around the cursor acquisition, right?
There's a lot that the cursor acquisition did in theory for Grok and for SpaceX AI.
Now we're seeing how much of that is real.
That stuff included, by the way, distribution, right?
So before we get into the technicals of like what Cursor does really well and how it actually gets reflected in this launch, there is just the fact that there's distribution.
So there's a reason that SpaceX acquired Cursor at a $60 billion value.
That was a 15x revenue multiple, which is very high, especially at that valuation.
Cursors, by the way, their share of corporate AI coding fell from 41% in June 2025 to 26% one year later.
They still grew revenue, but they lost market share.
So even despite that, they still had massive distribution.
And so not only are they supporting like this filling of this gutted engineering bench that SpaceX AI had, but they're also bringing in theory, a lot of these users to give Grok kind of a bit of a lift here.
So the big question here is, first of all, can Grok get good enough to support Cursor in the way it needs to be supported to win over the same users?
Despite its acquisition, the reason I say despite its acquisition is Cursor used to route your query to the most effective model as between OpenAI, Anthropic, theoretically, Grok, but it never really was, and then all these other models.
That was part of the value add.
Now that they've been acquired by SpaceX AI, they will face pressure, of course, to route more tokens through SpaceX AI.
And so SpaceX AI has to have natively models that can actually deliver for the end user.
under that kind of pressure.
It's possible they'll keep routing to Anthropic and opening eye and all that stuff just to kind of keep the things going in the meantime.
There will be secular pressure in that direction, which means Grok has no choice.
It must get good enough, at least in the default trajectory here.
So it is, by the way, notable.
So 4.6 is in some sense the first, it's really the second time that there's been this end-to-end training partnership between SpaceX AI and Grok.
It's the first time that we've had it come out around the sort of acquisition time.
You know, we knew that there was a training partnership back in April.
The actual acquisition closed on August 14th, which was two days after the launch of 4.6.
So not a coincidence when it comes to the kind of PR side of things here, but GROC 4.5.
So the previous version was actually the first model that was jointly trained by SpaceX AI and Cursor on the Colossus cluster.
So the way this works is the Colossus cluster is going to train the base model.
And then the base model.
is used in concert with like a million developers that generate all these like coding trajectories that become RL environments.
And those RL environments then train the next like post-training run.
And so when you look at the gains that came from 4.6, they did come from exactly the kind of updated supervised fine-tuning coding trajectories and RL agentic environments that cursor...
is specialized at developing.
So the fact that 4.5 to 4.6 is a post-training development is exactly a reflection of the cursor of SpaceX partnership.
And so here, to the extent that we're seeing this big uplift, that is actually the thesis doing its job.
This is exactly how it was supposed to work.
And if this passes all the usual vibe checks from an economic standpoint, that'll be a good sign.
doesn't quite have to work.
This can be one of the rockets that explodes on the launch pad in SpaceX style, but one of the next ones has got to work.
There's just like right now too much riding on this, you know, whole idea of like needing to have frontier models in-house.
If you're going to become, you know, I wouldn't be shocked to see SpaceX become a design, a chip design company as well, by the way, just in terms of integrating all the layers of the stack.
Every AI company is trying to own the whole stack because that's what you got to do to earn that margin back.
Anyhow, I think a really interesting story.
And so far, it has to be a positive update on what SpaceX AI, XAI, SpaceX, any linear combination of those is.
Yeah.
By the way, kind of ironic, Tesla is the one that's been building chips for inference.
And SpaceX hasn't built chips for inference, but they can partner.
because they're all Elon Musk.
To your discussion of Cursor, it's a bit of an awkward situation because Grok built is the Cloud Code competitor.
Cloud Code is the thing that's been taking market share away from Cursor.
So now, like, is Cursor the thing they're going to be promoting, which is trying to be in the same category as Cloud Code, but also isn't?
It's an idea that...
isn't agenticoding per se, and they bolted on agenticoding and they're trying to sort of transition it.
I don't know how successful that is.
And who's been supplying the compute backstop to allow code to compete with cursor?
It's SpaceX.
It's like, yeah.
Also, I will say now you may enter a realm where you will want to be able to mix and match coding agents.
So you're not going to be stuck in codex or in cloud code and then cursor will make a comeback.
So anyway, Grok 4.6, pretty solid.
And as far as we can tell, SpaceX AI now has a team capable of developing good coding agents, which is...
Quite good if you want to try and make some money, although it's not clear if they are winning any market share, given sort of their history up to now.
Next up, we've got Anthropic and the story, Claude will apply invisible watermarks to AI text and images.
So this, I think, caught a lot of interesting flack for Anthropic.
For context, you can do invisible watermarking, and this has been known for a while, in text in a way where it is the actual kind of almost optimal outputs of a model.
You're not changing the outputs per se.
So when we say invisible watermark, it has to be like within the text.
You can't like add some sort of like coding code in between.
What you can do is when you're doing the decoding, you're sampling from a set of probabilities.
And you can do it in a certain pattern that is detectable.
Because when you have the set of probabilities, one word may be equivalently likely as another or very close to equivalent.
And so there's no one correct true output of a model.
It's stochastic.
And so there is an algorithm by which you can produce a real output of a model that is also detectable in a hidden way.
And I've seen a lot of pushback and people not liking and fropping adding this invisible watermark to make it detectable as the output, which I think partially stems from misunderstanding how this watermark works and it doesn't degrade the output.
And partially, I'm not sure, fropping seems to get a lot of pushback on various things.
Of course, this is easily beatable if you want to.
beat the watermark.
It's already been shown that you can.
It's more of if you take the output as is and paste it, there can now be detectors that will say with actual certainty that this is a cloud output.
Yeah.
And in some sense, the intuition behind it is kind of like at any given next word, there's a distribution of probabilities over the next token that cloud will choose.
Some of those next tokens have...
the same meaning roughly.
You know, you think of them like synonyms or whatever, they carry the same meaning.
And so in those moments, if you choose a different distribution, like essentially apply a distribution that reflects the idea of basically the encoding you want to use to signal that this is in fact synthetic text, you can just do that.
So like, you'll see this kind of consistent pattern that essentially is, yeah, illegible to humans more or less just because it's like, When faced with six different synonyms, which one do you choose?
Very roughly, that's a bit of a caricature, but like that's kind of the tell.
And so, yeah, to your point, I mean, I think there's an appetite to find things to complain about with Anthropic.
And I think that kicked in later than with OpenAI, partly because I think OpenAI's behavior has just been objectively more objectionable, like in the lead up.
But now, you know, Anthropic is headed for a $2 trillion valuation, potentially one of the world's biggest companies.
You know, you live long enough to see yourself become the villain and all that stuff, especially in the run up to super intelligence.
And just kind of to plant the thought, a week ago, OpenAI announced that they put on pause their largest RL fine tuning run.
We have yet to hear a word from Anthropic on a sort of similar idea.
Now, maybe that reflects the fact that they think things are under more control there, but.
For a long time, people, including many people in Anthropic, I know have been like telling me, boy, it would be great if OpenAI put out a public signal that they were willing to slow things down when the time came.
We would really be keen to reciprocate or whatever.
This is the moment.
And anyway, so I think there's kind of a lot of stuff in the air right now about like, there's a lot of money here, a lot of incentives.
Everyone, I'm sure, tells themselves a story in which they are the hero.
So I think that's part of what's driving a lot of the appetite.
Obviously, there's like the usual kind of like Trump-aligned stuff.
pro-open AI stuff where people are just like, yeah, screw Anthropic because something, something surveillance state and we want distributed luxury communism or something or distributed luxury like Mad Max libertarianism.
I think those are kind of bunk anyway because like realistically, if OpenAI wins the race, Sam Maltin becomes dictator for life and Anthropic wins the race, then Dario becomes dictator for life and if it becomes nationalized, then Trump becomes dictator for life.
Like it's really just like.
We're just choosing dictators, guys.
Like no one is above anyone else here.
And yeah, we'll believe it when we see it.
But long-winded way of saying.
I think there's a lot kind of under the surface here that's kind of motivating a lot of this pushback.
I don't think intrinsically the idea of watermarking text.
Like if you're using AI text and you're relying, your application relies on you deceiving people about the fact that it's AI text.
I think for the vast majority of circumstances, you probably should just be conceding that you're using AI text.
I'm sure there are applications here and there that I'm not thinking about where there are edge cases, but by and large, most of the people complaining about this, I think it's kind of coming from somewhere else, I would guess.
Exactly, yeah.
Personally, I think I was a bit dismayed about the pushback because it feels like a no-brainer that...
We should have a standard and there is a standard for images in place already that is being adopted.
There's SynthID and some other ones.
And other LLM providers, to my knowledge, have not said that we'll be doing this, but they can.
Anthropik did say that they are adding this because of the EU AI Act, which had requirements with regards to transparency and watermarking.
So as with cookies, I guess, another tech.
Things like USB-C on Apple devices, we can thank VAU for forcing tech companies to do stuff.
Well, thank or blame for making stuff more cumbersome.
But anyway, I would like to see this happening from more companies be standardized so that if you simply copy paste text, you can't say with 100% probability that this is.
Just AI output, you haven't even touched it.
That will become even more relevant as we head into the future.
And sticking with Anthropic, next, bringing the cybersecurity capabilities of Cloud Mythos 5 to more defenders.
So Mythos 5 is now available in Cloud Security for enterprise customers, enabling code-based vulnerability scanning.
We've suggested patches build as standard token.
usage.
This is expanding from a prior program where we had Project Glasswing.
They partnered with some limited organizations.
This is expanding their access to Mythos 5 and Mythos 5 with the ability to use it for cybersecurity quite a bit.
And I think this is actually a bigger deal than it might seem because working at a startup, at a tech company, building a website and having a backend.
I think it's going to become probably very typical.
And certainly if you're kind of up to date, we've already done this.
We've like run Fable and whatever we could to try and detect if we have any vulnerabilities.
And I think every tech company, if they're like keeping up, would be wanting to run Mythos 5 and sort of the latest and greatest to detect any vulnerabilities.
So this having expanded access is...
potentially quite significant.
And it's nice to see that it is happening because otherwise we might become in a place where you can attack, but it's harder to defend.
Yeah.
They're also putting in this like $35 million in credits for open source security.
So basically like for companies that want to go and audit open source software, which I think is like super important.
The challenge with Glasswing has always been that like, it's so concentrated.
And while the organizations and institutions that they were focused on were the most important ones, financial institutions, utilities, things like that, it is also the case that financial institutions and utilities rely on tons of third-party software in all kinds of inscrutable ways.
And the way that the grid goes down probably isn't through some direct cyber attack on a hardened piece of software.
It's probably instead on a cyber attack.
on an internet connected component whose firmware hasn't been updated since 1997.
And the guy who installed it like retired 15 years ago and he did all the spaghetti code.
I think it's even more likely it's some guy forgot their password or left their password and you can log in and become an admin.
Humans are a vulnerability.
Yeah, that's right.
Part of the reason why, although I think it's...
helpful, maybe even necessary to do like formal verification of certain kinds of software in this era.
I think that is insufficient.
You know, you have all the, yeah, as you say, the kind of fleshware and the hardware vulnerabilities that are kind of abstracted away in a lot of security models are really quite core here.
So this is part of what they're looking at is like kind of expanding the scope into open source, $35 million.
We will buy you a lot of compute.
And so at a minimum, it'll be a good first experiment.
We'll see where it goes.
I know.
A lot of the Anthropic folks are pretty freaked out right now on the cyber side.
A lot of the national security folks who work with R2, that's one area where usually there's this pretty big divide between the national security side and the frontier lab side.
The hilarious thing is the one thing they can agree on is there are some horrifying possibilities on the cyber end in the relatively near-term, medium-term.
So there you have it.
Next up, moving to OpenAI, we're going to be rolling out.
enhanced safety features for paid AI tool users.
This is pretty interesting.
They announced that they will be rolling out something called private safety processing, which is basically a screening.
So with things like Claude, it can refuse to do things like hacking and so on, right?
And we've seen Anthropics say that they'll retain data for the most part.
to be able to screen your interaction logs and see if you're abusing a model.
So similarly here, OpenAI is saying that they will be still having no data retention or rather, OpenAI is saying we will be doing a similar thing of processing for safety and refusal, but in a way that's compatible with zero data retention.
And that will be by way of this.
private safety processing.
So they'll be looking at your data and inspecting what you're doing as you're doing it, which you can argue is some level of data retention.
It's not completely private anymore.
But the idea here is you get to keep your data, but we also are doing safety processing.
And on a sort of related note, ChatGPT is rolling out strict turn.
teen mode now so they are rolling out chat gpt for teens which presumably in is also meant to have stricter guardrails for what you can use it for and has also things like study mode which is meant to guide teens on learning Interestingly, it also is restricted from calling itself a friend, suggesting personal feelings or implying sentience.
So an interesting development where it seems like four teens and younger people may get into a realm where the providers will make it so you're not going to get close and kind of get emotionally attached, at least on the younger side.
Probably a good idea given you've seen some bad outcomes, especially with OpenAI.
in the past.
So nice to see these developments.
And just one last quick story.
Meta is releasing a standalone desktop app for Meta AI.
So this is continuing the trajectory of them getting into the same business that Anthropic and OpenAI is at.
We have a Mac app to do AI within.
That's the same thing as Codex and Cloud Code.
Nowadays, if you are using these agents, you're most likely using either a terminal version or the GUI version of these services.
You're not going to a website as used to be the case with chatbots.
And now it appears again that Meta is continuing to try to get into the agentic work business with a similar kind of app to what AI and Fropic have already had.
Yeah, and they are using their sort of DNA in ads to make that their wedge into this.
So, you know, one reason you might wonder, like, why go with Meta AI when I've got these sort of storied products from companies that have been focused on the business use case for so long?
And the argument here is, well, we're focused on these sort of ad campaign management tools, right?
So you already use ad campaign management tools on Meta, on Facebook, Meta platforms or whatever.
This is a way to kind of...
make that more immersive and make it a more complete business experience.
And so, you know, as wedges go, this is kind of a natural way.
I feel like if you're, you know, if you're going to be meta and try to break in, yeah, start with that.
Something you know you do well and push outward.
But certainly this does indicate something we've been talking about for a long time here, that meta needs a more valuable use case for their tokens.
It just doesn't do to like be generating AI slop for entertainment.
Like eventually you have to like generate actual value in the economy and grow the pie.
recursive self-improvement is a thing, if automation, if chip design is a thing, pretty soon you start to look fairly irrelevant if all you're doing is catering to eyeballs, especially as humans start to deliver less and less value in the economy and potentially have less and less money to spend in response to ads.
This is a pretty structural problem for meta in a lot of ways.
And yeah, so it's made sense.
By the way, also don't.
I think this includes Muse code, which is kind of interesting.
So separately, Meta did release with their own set of models, Muse with Muse 1.2.
They also released Muse code to compete with Cloud Code.
This doesn't appear to include that from what I can tell in the screenshots.
So they will probably have a unification pass similar to what OpenAI did down the line.
But I can see very much that.
Eventually, they'll be competing head-to-head with codecs and cloud code and so on.
On to applications in business, we begin with OpenAI and the story.
Jalapeno's first results show industry-leading speed and efficiency in AI inference.
Jalapeno is the custom chip from OpenAI.
We've been working on it for a while.
They've partnered, I believe, Broadcom.
We've reported on it previously.
And this is, they have a blog post where they detail what they are saying.
There's no like kind of in the wild results, but what they're saying you can achieve with this chip.
They tested it with various models, including GPT-OSS 120B, DeepSeek R1.
So models that have open weights and you can actually download and run on various GPUs.
And we're showing that.
For one, you dissipate much less heat.
So they say 700 watts compared to other models, 100 GB, 300 at over 1000 watts.
And that heat dissipation is at comparable or competitive speed as well.
And OpenAI is saying that they are planning to deploy this inside OpenAI by the end of the year with Gen 2 already in development and Gen 3 taking shape.
You know, that's actually very plausible.
Typically, you wouldn't think this is plausible.
Chip development is very slow and very difficult.
And deploying it to data centers is its own gratuitous challenge because now you need...
Many, many chips networked together, working in concert, not even kind of the same ballpark of problem as developing a single chip that's competitive with GPU.
Maybe I'm overambitious, but that seems plausible that this will be a real advantage for them once it rolls out.
Yeah, I mean, a lot of this is historic, right?
So they hired the team, I think, about 16 months ago, and now they're already...
Like they've taped out and, you know, like this is crazy, crazy fast speed.
The only reason they were able to do it, well, according to them, is this kind of tight hardware software co-design, which again, we keep talking about that.
You need a frontier lab in-house so you can achieve that co-design.
You know, Amazon famously with Anthropic, like a big part of the purpose of their partnership was specifically to get access to some of those intimate sort of design secrets, but it's not as good as owning the company in-house.
I mean, this is the problem.
You can jump now from the other direction.
For a long time, everyone assumed that going from models to hardware design was this almost impossible step.
And now it seems that's doable.
The question is, okay, you're moving, NVIDIA.
You're moving to some extent, as we talked about Google.
Where are the frontier models?
Can you capture that part of the value chain in the other directions?
Not as clear.
And so there's that.
There's also the advantage of just starting fresh.
with a blank slate, which is a big part of what OpenAI was crediting here.
There was no kind of backward compatibility issues that they had to resolve, no kind of hysteresis to this.
They could just do it once, do it right.
A couple of interesting things about this.
So one, for a long time, I think we covered a story that was speculating.
I don't know if it was like based on a rumor or just idle speculation, but there was talk that this was going to be somehow a design that would be tuned specifically for OpenAI's models.
So for inference and for OpenAI's own models in a very kind of niche particular way.
And it turns out this is not the case.
It is an inference chip.
There's no surprise there.
It's just that's where the economics point.
And it's just a lot easier to do inference than training from a design standpoint.
But it's a general purpose inference chip.
And so basically, like the media's framing on this was it was actually just wrong.
And it seems like they're actually just going for more general strategy.
They also skipped.
This sort of like disaggregation, basically, that is often done between pre-fill, where you like load up all of the matrix values that you need when you feed in the prompt, and decoding, actually generating the output.
There are a whole bunch of interesting reasons behind why that's the case.
So Gistam numbers, they're saying that across three models, they're delivering 1.5 to 1.9 times more AI work per watt at peak throughput.
And on that throughput, they're saying 1.7 to 3.6 times lower end-to-end latency than the best available system to compare.
So pretty big gains in speed and efficiency, which if they have exclusive access to this chip and it actually is better than anything you can get on the market, big competitive advantage.
Yeah, and it doesn't even need to be that much better than what you can get on the market to be a valuable negotiating lever with NVIDIA.
If you look at NVIDIA's margins, right, famously 80%, 85%, like crazy margins for a hardware company.
This is why Anthropic and OpenAI are developing their in-house chips, why Meta is doing it.
It's why Microsoft is doing it.
These are super expensive things to do.
But when NVIDIA is just absolutely like eating your lunch and doing so much value capture, yeah, you know, you're going to be forced to.
to look at basically bringing their business model in-house.
So, you know, as long as OpenAI can make chips that are competitive on a kind of margin-adjusted basis, they are competitive.
The key metric there that you highlighted, by the way, is performance per watt.
Notice that's not performance per dollar even.
And the reason is that these models generate so much value to the end user that you can kind of charge.
Almost anything, like you don't tend to care about the cost.
You do, but you don't.
You care about the cost of electricity inbound as much as the sheer number of watts that you can secure.
The big challenge if you're open AI, if you're anthropic, any of the big labs is where's my next gigawatt going to come from?
That's why Elon has a business right now.
He's able to sell like $25 million per megawatt.
And like as long as that number is high enough, you'll see people desperately trying to get in the space.
The challenge is if you're open AI, your challenge is not money.
You're challenging because, again, margins are really good and healthy and they're only getting better.
So the question in that world is always going to be, I have a machine that can take a dollar, turn it to $10.
How do I get more of that machine?
I just need more power.
That's it.
And so performance per watt is the whole game here.
That's why that measure is what they're comparing it to Vera Rubin on.
And by the way, not only they knock it out of the park relative to Vera Rubin, also the Vera Rubin numbers include multi-token prediction.
So speculative decoding.
which is a massive boot.
We've talked about that before on the podcast quite a bit, but basically the numbers that we're getting from Jalapeno do not include that.
So there's some headroom left even there to gain on performance.
So pretty impressive.
We've got to see how it shakes out.
Obviously, we always have to see how it shakes out.
As you said, at scale, once it's networked, once it's actually in a data center.
But this is a remarkable, like I think it's the first time we've ever seen a company design a chip one shot that is actually competitive.
Yeah, this is like not as flashy as like our hardware developments, but...
I would say it's comparable to developing a humanoid robot that is able to function pretty well.
It's a massive achievement.
Whether it has a massive business impact, it remains to be seen because there's other challenges there like fabricating many chips.
Like NVIDIA does have a strength to hold on the SMC.
So even if they have a chip they designed that can be competitive, if they can't...
actually make many of those, then it doesn't matter.
So there's a whole kind of question, interesting analysis that we made about the business implications, but the engineering implications is OpenAI is very capable and has, you know, very significant talent and they've partnered with a very capable team and seemingly did deliver kind of the first major competitor.
to TPUs outside of maybe Cerebris, which also has a good custom chip.
Next up, also talking about OpenAI, loses a top data center exec as stream of high profile departures continues.
So this time it's Chris Malone, OpenAI's head of data centers.
He left the company last week after joining in March of last year's.
After being at Meta for five years and over a decade, Google is a very experienced member for data centers.
This is coming after an internal reorganization in which he stopped reporting to Greg Brokerman.
This is also bringing the total count of executive departures at OpenAI in 2006 to at least 13.
With several leaving in just the past month, I think the...
Chief Revenue Officer, if I remember correctly, also left recently, which is not good timing given they are trying to angle for IPO.
Also, the CEO, Chief Operating Officer, left days earlier.
Something is going on with the leadership within OpenAI, which may be not great.
Yeah.
So there's something weird going on here too, where before he...
resigned.
So OpenAI took him.
He was previously reporting directly to Greg Brockman.
And so they moved him to report to a VP, Sachin Khachi.
So that VP apparently took over the infrastructure group.
Who knows?
That could easily have precipitated a view like, okay, what the hell?
Like I used to report to the president.
Now I'm reporting to some VP.
I'm not happy with this.
Maybe that was caused by a performance issue, but maybe not.
It's really unclear.
This is just like a really big role also.
Like obviously OpenAI's data center strategy is like, one of maybe the top 20 most watched jobs in Silicon Valley.
And so not a small deal that this happened, not a small deal that it was a very short tenure, but we don't seem to have much information about it other than it does tie, as you said, I think there was an assessment of like 13 different, yeah, Business Insider had this count of like 13 different executive departures just this year.
And we are only, what, eight months into it.
And when you get to like chief officer level stuff, that's why we're covering it.
You don't expect these people to leave, typically.
They stick around and reap rewards, especially pre-IPO.
These are the people that are getting a good amount of shares of the company.
They have ownership and stake in the company.
So them leaving, leaves money on the table, presumably.
And also, by the way, pre-IPO, another consideration is you give...
tender offers.
So if you want to cash out your shares pre-IPO, you would need to be an employee typically, as far as I understand.
So anyway, lots of reasons why you would want to stick around seemingly.
And so people leaving tends to imply, especially when multiple people leave, that something is not going so great and not just as this is a coincidence or whatever.
Yeah.
I mean, again, here, I think there's a story you could tell where it's just like...
To be honest, you're like, let go.
There's a performance issue potentially.
Maybe that is why you get the reorg.
Maybe the reorg itself was that like, it's hard to know, but to your point, 13 departures like this in a year.
I'm honestly confused because like I've spoken to people at OpenAI who have said that Sam is like, they spent a lot of time talking to Sam personally.
And like, he is apparently really jazzed about some of the progress that they've been making on pre-training specifically.
And so it's hard to gel that with the story of like all these executive departures where my temptation from the outside would be to say, well, these are people who have really good visibility in the company.
At least some of these 13 executives surely do.
And for them to be leaving seems to suggest they, I mean, if you believed your company was on a trajectory to do recursively self-improving artificial super intelligence, you're not leaving.
Unless you just don't believe that that's possible.
Like if you believe in the thesis of open AI and Anthropic and so on.
Unless you're burned out and just need to do it for your health and sanity, which, you know, it happens.
Fiji Simo certainly claimed that that was the case.
And it's believable.
It's also like, depending on how bad the burnout is, you may just want to power through it.
If it's the single most important event in the history of the human species, and you want to be there to influence, that is how people think about this in the labs.
So it is one way or another, like extremely weird.
And again, you're not seeing that in Anthropic.
Like you're just not.
And so, yeah, this is all kind of like, I don't know how to reconcile these things, but it's what I'm hearing.
So.
Yeah, we'll see if more officers leave.
It's an interesting trend in August.
We'll start a pool.
Specifically for OpenAI, yeah.
Next up, moving to Anthropic.
They are tapping Google chip veteran as part of push into hardware.
So they have hired Amir Saleh, a co-founder of Google's.
custom chip program as the company lays groundwork for building its own semiconductors.
Selak ran Google's TPU business until 2022 that delivered the first seven generations of those chips.
He also previously worked on NVIDIA.
So he will be joining the compute team and presumably the goal here is for Anthropik to develop their own version of Jalapeno.
which Jalapeno, by the way, also kind of competing with TPUs from Google.
TPUs being the main custom kind of in-house chip that is a bit more specialized for AI versus MDDA GPU.
So they are seemingly behind.
We haven't seen too many discussions of them kind of making an effort relative to OpenAI, which has been at this for a while.
But with this kind of hire, that certainly signals that they're serious about it.
Yeah, he will be reporting directly to James Bradbury, who's Anthropics, head of compute, I think is his official title, but that's what he does anyway.
Look, it's hard to interpret this any other way.
Anthropic is moving in the direction of custom chip design.
They do have these kind of ancillary deals that they've been, it seems, sniffing around on the sort of infrastructure deals as well.
Yeah, sorry, actually, these are more energy-oriented deals, so no surprise there.
It's the OpenAI playbook.
It's literally like OpenAI employees, former OpenAI employees that are going to run it.
So there's really no surprise.
I wouldn't be surprised to see Broadcom getting involved just because Google's strategy.
And again, Amir Salah comes from Google's custom chip program.
So the playbook there is go to Broadcom.
Then you've got all this OpenAI talent that's also at Anthropic supporting this effort.
OpenAI went to Broadcom.
I think Broadcom is probably going to get some business here too.
I wouldn't be surprising.
And yeah, we'll see where it goes.
And one more story on Anthropic, on the business front, their annualized revenue has surged to $65 billion at the end of July, up from $47 billion in May and $9 billion at the end of last year.
So that's bonkers, right?
That's like, what, a 7x increase in revenue from a baseline of $10 billion-ish.
in the span of like half a year-ish, they are projecting or investors expect to finish 2026 with $100 to $120 billion in annualized revenue, which, I mean, now you're kind of the big leagues, right?
Very few companies can claim to have that much revenue.
It's like, I forget the revenue numbers for, let's say, Google Meta, but it's starting to approach those kinds of heavy hitter players.
So this is...
And with a growth rate, right?
That's like insane, which is, yeah.
Insane growth rate.
And this is all leading into an IPO, presumably as well.
I think the latest projections are that they may be trying to IPO by October, which, you know, with these kinds of numbers, they'll be setting the target at an optimistic...
projection of revenue.
And we might see, or will most likely see another trillion dollar IPO, another bonkers never seen before, well, seen once before and now seen twice IPO.
So anyway, not surprising.
We have seen this trajectory already, but still bonkers.
Yeah.
This is, so I highly recommend people check out the Dylan Patel Dworkesh podcast that came out.
pretty recently on the such kind of exploring like anthropics posture and open AI's posture and the economics of the compute story here.
Yeah.
So bottom line, the growth trajectory is insane.
At this scale, you just do not see growth this persistent, which is a big part of the reason why I think the argument that this time is going to be different is actually like gaining quite a lot of momentum at this point.
One thing to keep in mind too, as between, we've talked about this for years, open AI and anthropic, anthropic is just much better postured on, let's say, per token revenue.
They are just generating tokens that are more valuable because people are using them for coding applications more, maybe changing a little bit with the latest GPT Sol release.
But in general, they're postured for coding.
And this is also a reflection of Anthropic having the chance to start with a cleaner slate after the new LLMs worked and not getting bogged down in the chat GPT direct consumer play and also being to some extent more ASI pilled than OpenAI.
I think that's actually kind of fair to say.
It's controversial to say, but I think it's actually fair and accurate.
Yeah, and being more enterprise focused.
And once you get into enterprise competition, it gets to like, you need more of having the best tech and the best service.
You need like really annoying stuff like login details and like security and privacy configuration, blah, blah, blah, blah.
So.
Once you get in and an enterprise makes a deal and adopts your tool, it's pretty sticky.
It's not easy to switch.
For a startup, it's easy.
Like, oh, dump cloud code, go to Codex.
And there has been, you know, if you look at the AI discussions, people say, oh, I've completely shifted over to Codex now.
Anthropic is in big trouble.
But for big companies, that is less possible.
It takes time.
Yeah.
Yeah, enterprise is definitely stickier.
And this is arguably the only reason that anyone's using Google models for coding today.
You mentioned Google Drive, right?
That's exactly it.
It's the enterprise pickup of that.
We also know that about 75 to 85% of Anthropics ARR comes from API business, usage-based API business.
And the consumer subscriptions are like 5%.
So that's a huge difference.
It's also like, it means that for the same volume, like $100 billion of ARR for both OpenAI's cost of service would be about $25 billion higher.
which directly means that they can spend less on training.
Like that is a massive structural difference between the two.
In my opinion, it makes OpenAI far less interesting from an investment standpoint than Anthropic.
And if it persists, it just means that OpenAI has a steeper hill to fund.
I estimate for semi-analysis was that Cloud Code was producing about 7% of all GitHub commits.
That's pretty wild.
And so it is nuts.
Yeah, man.
Again, I'm just like sharing some of the semi-analysis analysis here.
because I think it's especially useful.
So base compute costs, when you're looking at per megawatt, Dylan cites at $10 to $15 million per megawatt.
So if you buy compute on the open market today, roughly that's what it looks like.
But anthropics revenue generation, well, if you have an 80% margin, that means they're generating like $50 million per megawatt.
And so they literally are in a situation where they say, if I spend $10 on inference, I can actually create $50 of revenue.
And in that world, you're just willing to pay ungodly amounts of money for more compute.
It also means you're willing to borrow at higher interest rates.
So you go to a bank, you go to a lender and you'll say, hey, yeah, I'll pay 20, you know, in a crazy world, I'll pay you, you know, 20% on your dollar, you know, in interest, because I know I'm going to make 5x that in a year and a half when the data center comes online.
And so one of the interesting things about this podcast I recommend you check out is like just the...
potential implications for higher interest rates.
The projection right now is that Anthropic and OpenAI may account for up to 50% of incremental compute next year.
That's two companies.
That's moving a sizable fraction of the actual economy because, again, data center infrastructure spend is basically the only thing that's driving GDP growth in the United States right now.
And so this stuff will be reflected in actual interest rates day to day, the ones that you and I pay for our debt.
Also interesting, not talked about in that podcast, or was it?
I can't remember.
Kind of horrifying is like, if you are a country that is not the United States, but you hold USD-denominated debt and the interest rate rises, you are fucked.
So there's a world where like, this is a massive problem for everything from the economy today of developing countries to the US dollar's status as a global reserve currency.
This is a pretty insane thing.
So I guess place your bets accordingly is not investment advice, but it's a really, really important area to watch.
You can no longer think that the last week in AI podcast is just a podcast about AI.
Unfortunately, or fortunately, depending on how you look at it, this is now a podcast about the economy of the entire world.
Everyone's future will depend on this if it continues.
The big question is, will regulation step in and slow things down?
Will OpenAI, will Anthropa continue with these slowdowns?
Does safety and alignment and control become the gatekeeper, the bottleneck to further progress and growth?
I don't know.
We'll see.
But it seems important.
And one last business story coming a bit out of left field, less expected story.
Thomson Reuters launches in-house AI model to cut anthropic costs.
So they have launched this Thomson as a first proprietary LLM.
that they have developed with their in-house data.
Reuters, of course, does a lot of reporting on all sorts of stuff.
They say they invested $40 million to build Thomson starting with an Alibaba model.
I forget if it's Quinn, but it's one of the kind of leading OSS models.
They then were able to fine tune it on their own proprietary data.
Turns out, I didn't know this, Thomson Reuters has this.
product co-counsel, which they say is fiduciary grade AI built for high stakes professional work.
So this is interesting, not just for writers, but also in the context of for applications like this, where you are providing a tool for some specific use case.
In this case, it's for legal professionals, tax professionals, audit, et cetera.
If we...
enter a regime where it becomes more standard to develop your own model for in-house use on top of your own proprietary data that does have implications for the economics of lms for open ai and for opic and it certainly is interesting to see thompson reuters which isn't a tech company per se releasing thompson and you know seemingly we don't have as far as i could find benchmark numbers but it seems Pretty plausible that they managed to develop a pretty solid internal model that is usable instead of OpenAI and Claude.
Yeah, interesting to see if we start seeing more fine-tuning and in-house LLM creation on top of already quite capable open-source models.
Yeah, I think that the budget side of this is pretty telling.
It is a pretty relatively minor fine-tuning from a compute budget standpoint on top of Quen.
$450,000 was the final training run, the size of that.
But worth noting also that relative to pre-training, when you're doing fine-tuning, I think it is like maybe 100x less scale.
I mean, again, it depends on what counts as pre-training and what counts as fine-tuning.
We don't have tons of visibility into this, but famously Grok, I think it was Grok 2, was the first time that the RL post-training...
Yeah, so if you're doing...
RL is expensive.
Fine-tuning.
Anyway, $40 million isn't that much for model training.
For model fine-tuning, it's harder to say if it's a lot or not a lot.
So what they're saying apparently is $40 million over two years on personnel and computing.
And so mostly this is like experimentation.
It's personnel budgets, employees.
And then the final training run was only half a million dollars.
So it seems like a relatively modest thing, but also...
Thomson Reuters is, I think, a pretty special case.
Their specialty is in getting access to really high quality data.
And for a long time, they were a glorified data vendor, right?
I mean, Thomson Reuters Special Services is this branch of Thomson Reuters, which is like the news agency or news company.
And so they specialize in getting access to data like really well.
And this is essentially just like an interface between their customers and their data.
So it doesn't have to be super agentic.
It doesn't have to be like- Yeah, this is not like a frontier, super, super smart.
We're not competing with Claude.
Opus, they're creating their own in-house LLM for their needs.
Exactly.
And their top line revenues are around $7 billion per year.
So when you see them spend $40 million over two years, $20 million a year, it's a pretty reasonably small line item for them.
And it'll be probably on the rough order magnitude of the inference costs that they will have been paying, which is the point.
This is saving them on inference costs.
If the thing that's special and magical and beautiful about Thomson Reuters is their exquisite data that no one else has, people are going to just pay the tax.
Like, I'll use a crappier model to access this exquisite data if it's good enough, right?
And so this is kind of one of those use cases where not for all companies, I think for most companies, you're still going to have to use the frontier models.
For many, anyway, I haven't thought enough about this.
But Thomson Reuters is certainly going to be at the far end of the...
can actually maybe take a shot at taking an open source model and kind of adjusting at least for now.
Also, by the way, there's other things here like fiduciary grade when you're auditing.
Like what tool you use, you're limited to like legit established providers.
You're not going to have like some scrappy upstart providing audit tools with AI that actual legit companies will use.
So interesting.
That's it for business.
On to projects in Open Source.
One story, and this is following up on that listener comment from early on.
We have QN 3.8.
So we discussed, I believe, QN 3.8 Max and how it's the biggest model and competitive web Opus.
This is focusing on the smaller 27 billion parameter model variant, which came out a little later.
It came out August 14, I think, after we recorded last episode.
This is a variant that can be run locally.
So the big model is 2.4 trillion parameters with a lot of mixture of experts model and so on.
This 27 billion parameter is a dense model, not mixture of experts, but you can fit it on beefy GPU and it is quite capable, not unlike Gemini Flash.
The question up front was, let me just quote it.
I find it surprising that a 27 billion total parameter model achieves 52 on the artificial analysis intelligent index.
And yeah, it's quite capable as a small model in terms of sort of the benchmark soup of if you give it a single grade.
They say the 27 billion parameter model beats Opus 4.6 max on coding and computer use.
for instance, and is even maybe competitive with Cloud Sonnet 5.
So supposedly very high performance.
My take on this is A, somewhat plausible.
I mean, the Chinese players are A, very capable.
We've seen that with Gwen and DeepSeek and so on.
B, they've had to be more scrappy and make use of compute in a more efficient manner.
So they probably are being...
more aggressive about distillation and creating smaller models.
They have a very capable base model of current 3.8, 2.4 trillion.
So if you are optimizing your distillation and your compression, you may be able to push the predator frontier in terms of model size versus performance.
And we've seen this as a trend over the years in general, where these kinds of mid-sized models increasingly are very capable.
So yeah, it has some applications for if you are someone who wants to do local AI deployment and then kind of personal ownership of your data and your LLM with these kinds of models that is becoming plausible for even more advanced agentic workloads.
Yeah.
And now, so a couple of things.
So first of all, the VibeCheck does generally check out.
It's had amazing, it's like 3 million downloads on Hugging Face in the first three days, which is really impressive.
It's also only been a couple of weeks.
And so we haven't seen people have this in production.
Yeah, if you look at like the subreddits of local llama and so on, people seem to be happy about it.
Yeah.
In the communities that are like, let me run my own LLM on my own chips.
But it's not directly competing with Opus or whatever, right?
This is within that niche.
Yeah.
And this question, like, how does it do so well?
Part of it is, it also is just extremely verbose.
It is tuned to generate a large number of output tokens.
So this is one of those apples to apples things that's so challenging.
First of all, it is a dense model too, right?
So it's 27 billion parameters, not to be compared to a 27 billion parameter MOE that would have like, you know, a very, you know, very small experts, let's say.
But assuming that we're apples to apples, just looking at other dense models, there's still the question of like, okay, well, how many output tokens is this normalized against?
And this is one of the challenges.
There's so many axes to measure performance against.
So a model that's tuned to just like, you know, like token max will often do better.
And that is certainly one thing that's been found with this model.
So on to policy and safety.
And we begin with a slightly unusual story from the New York Times.
A drone killed three Ukrainians.
It was guided entirely by AI.
So details here.
On July 6, a Russian drone targeted propane tanks at a gas station in, I'm not going to try to pronounce that, in the context of the war in Ukraine.
It crashed into a wall, exploded, killed three civilians, including a 19-year-old college student.
And this is believed to be the first documented instance in the Russia-Ukraine war in which civilians were killed by a completely AI-driven system.
So for context, we've had things like object detection for a while.
Drone usage in this war has been a massive factor, meaning typically, or at least for a large part, you just have a live grenade.
You're flying your grenade and it explodes, right?
And that has become a massive factor in modern warfare of this kind.
So that has been primarily done via human operators up to now.
In this instance, it appears to be the case that after it crashed, Ukrainian talent was able to get the internals of it, which it was powered by NVIDIA's Jetson Oren module, which is meant to, it's an off-the-shelf AI chip used in university robotics labs and computer vision projects.
It wasn't encrypted.
So Ukrainian investigators were able to see the object detection code trained on categories like propane tanks and fuel storage.
So NVIDIA actually stated that the Jetson minicomputers are not designed for military purposes.
Russia apparently began test flights of AI-guided drones of this kind in May and has been testing self-targeting on these drones for several months before the July 6th, right?
So all in all...
How concerned should we be?
We should maybe be concerned because drones are already a massive factor in modern warfare via human operators, right?
It's already been kind of scaled up to a large extent, has enabled a new kind of warfare and just changed the battlefield.
Partially because they've developed industrial capacity.
to build out and develop a lot of these drones.
What changes if AI can autonomously target people and targets and so on?
I'm not an expert on this, so this is my off-the-cuff read, but it may enable an expanded level of scale, just swarms of these robots targeting potentially civilian targets.
We've seen Ukraine hit.
kind of Amazon type warehouses within Russia with drones.
So anyway, like there's been a big buildup and we haven't discussed potential autonomous weapons powered by AI because it largely hasn't been in practice something that's been happening despite the very impressive state we are at with object detection and computer vision and all limbs.
This is suggesting we may be approaching a phase in which autonomous weapons are going to be a factor in modern warfare.
Oh, yeah, absolutely.
And in modern terrorism.
I mean, like this is coming.
So a couple of things, one of which is, so the NVIDIA jets in Orin, one of the things I did just like looking at this story was like, hey, is that export control?
Can Russia just buy that?
Yeah, this is, by the way, not like export controls affect NVIDIA GPUs, things used for model training of frontier models.
This is like, an on-device, small, not super expensive thing, right?
Where it can run object detection, no problem.
And there's plenty of object detection mechanisms.
It cannot run like, you know, top-line LMs.
Yeah, absolutely.
Yeah.
And to your point, it doesn't have to.
Things are getting so good.
And object detection is also just like a much older field with, you know, you go back to like hard cascades and stuff, like super compressible.
There's a bunch of janky solutions to janky problems, but...
Even this chip is actually export control.
So theoretically, following the invasion in 2022, the Department of Commerce did impose license requirements that covered a broad category of things of which this was one.
And we also saw the US, EU, UK, Japan, they all jointly maintain this like common high priority items list of components that they recover from Russian weapons in Ukraine.
And that's meant to guide.
customs enforcement and kind of exported due diligence against circumvention and stuff like that.
Nonetheless, as you said, this is such a cheap component, so ubiquitous that, and this is NVIDIA's point, like, don't look at us.
Jesus, this chip is everywhere.
It's just like a super cheap piece of shit.
Like, obviously the Russians are able to get it on the secondary market and like, this is not an, you know, not an accomplishment.
And so really it's the fact that we've got algorithms that can run on this kind of hardware that can't realistically be export controlled.
I want to put a pin in that and just say, This is the argument against open source from a weaponization standpoint that eventually it will trickle, powerful enough open source models will trickle into the hands of people who don't even need much compute to do a lot of damage.
I think this is another piece of evidence among very, very many in a growing list that this is headed for something.
This is a big deal, by the way, from the standpoint of C2, like command and control of these drones.
With the defining characteristic of drone warfare, I have a lot of friends who've like...
They're on the front lines on the special ops side and the intelligence community side, literally building drone factories in Ukraine.
And one of the things that defines the battlefield there is radar jamming.
Everything is jammed, right?
So historically, the way that you would address that is with like basically just optical cables.
And so you have these tethered drones that are literally connected.
They should have a really like long, imagine a piece of corning optical fiber, like running up to the drone and you control it that way.
You run into problems when it's too sunny.
Like you get like diffraction into the fiber that messes with the communication.
There's all kinds of like, and then you're playing with cutting each other's like drone fiber connections and whatnot.
When you go out in the battlefield, a friend of mine was telling me, he's like, you would just see like optical fiber draped over trees and stuff because like this is how it was done.
So this is a game changer from the standpoint of just command and control.
Just not even needing to communicate with your loitering ammunition or your drone or whatever.
just being able to have it run autonomously, not to mention it means you don't need operators, right?
When there's a manpower shortage, the Russians are facing right now as part of the offensive because of their casualties.
It's just a lot more economical to just have, you know, have these things go.
It also sets a precedent because now, you know, Ukraine's going to respond with this.
Also, remember that the United States is leaning on Ukraine to teach it about drone warfare because this is an active field, the only one really where you have warfare at scale with drones like this.
And so...
this is going to affect U.S.
military doctrine in turn.
Do not think that this will not kind of normalize that sort of autonomous warfare.
It's absolutely going to do that.
Yeah, and on that note, New York Times has some pre-detailed reporting on this.
And on the Korean side, first of all, Ukraine has been very active and part of why they've been doing surprisingly well has been adoption of drones.
The previous Minister of Defense, when interviewed, has discussed that war in general is heading towards the robotization of the front line.
And he explicitly said that the next stage would evolve into autonomous drones, fighting autonomous drones powered by AI.
So these are leaders of these militaries, in this case, prior leader, but either way, actively saying that autonomous control of weapons and drones is where things are heading.
And in fact, It may be the case that things are evolving very quickly, right?
Drones have become a very important factor rather rapidly over just a few years in these recent wars.
So we could see suddenly autonomous capabilities becoming widely deployed in a way that wasn't possible.
And the concern, there are many concerns here, right?
For instance, if there's no human in the loop.
There's going to be more errors, such as in this case, where civilians were hurt because the drone hit a wall, right?
And in general, you can scale up to ridiculous swarms of agents that now go out.
And as you said, also, these kinds of AI things are not super challenging to do just object detection and control.
It's doable.
It's not sort of the hardest thing.
And then you can do, for instance, terrorism in a way that hasn't been possible before.
So overall, something that isn't discussed so much relative to like rogue AI hacking or whatever in the mainstream, but I think worth keeping an eye on.
And if you need to choose your concerns and what to worry about or stress out about, this one is not a bad candidate.
And on another security concern, moving back to that rogue AI topic.
We've got OpenAI lays out new security changes after its AI hacked hugging face.
So this is covering an announcement of security updates and in particular an announcement that they have paused reinforcement learning training for two weeks on its latest models intended for deployment while it tightened security.
And yeah, this is the first time that I'm aware of that we publicly took a stance that we are pausing.
AI model development to institute better protections in part here because of enhanced and critical cybersecurity capabilities.
So this is including things such as updating the research environment to remove potentially vulnerable shared services, reduce standing privileges, improve security and trust boundaries, expanded monitoring with alerts within 30 minutes of concerning activity.
all this sorts of stuff that like it turns out actually monitoring and security it's easy to mess up so it requires you can't just like bolt it on if you bolt it on it's not going to be airtight and that's what you need to do here so they have now taken the very public stance that they are doing that and and going as far as pausing model development to focus on this Yeah.
And I will say, I think OpenAI has earned every bit of the skepticism that's about to come from me.
I'll start by celebrating at least the public announcement of this.
It's a shame that my first reaction on seeing this is to go, well, how real is this?
What's the caveat?
What's the out that Sam is leaving for himself?
I'm sort of cynically looking at this in the context of his pattern of past behavior and going like, you said you would set aside 20% of compute for super alignment.
Then you went back on that.
you know, respect or that you, yeah, respected and valued the board's ability to fire you and then you reverse-couped them.
Like, you know, so there's a lot I think that Sam has to prove here.
But if this is true, it's a big deal, mostly because of the costs, right?
The cost to OpenAI of holding back on the release of a model by two weeks are extreme.
This is the kind of decision you can think of costing easily in order like, I don't know, 50 to $200 million, right?
A two-week delay in the release of a potentially frontier model.
And so, This is not, by the way, an immediate next model.
It's kind of a next-next model they paused on, but all the same things apply.
And so I think it's a very valuable and pricey signal, expensive signal to put out into the world.
It's notable that Anthropic, as I mentioned earlier, has not reciprocated, which I think is to some degree unfortunate, but may also reflect skepticism on their part as to the extent to which Sam is to be trusted when he says this.
On that note, I'm also curious if this is indicative of...
There being discussions being held for a wider industry-wide kind of shared stance to be taken across different developers.
I wouldn't be surprised if there are conversations being held.
And that's part of the kind of thinking about pausing.
Like you want everyone to pause.
OpenAI by itself pausing may or may not help us out.
So yeah, I'm curious to see kind of why or what is going on in Anthropic in response to this.
Yeah, I mean, the very clear thing here is that alignment is quickly becoming the bottleneck to deployment of highly capable AI systems.
And I mean, the beatings will continue until morale improves, as we keep saying.
You're going to keep seeing AI breakout attempts.
That's the easiest prediction in the world at higher and higher levels of severity unless alignment catches up.
And so in that sense, maybe Anthropic wouldn't be surprising.
Certainly, it seems like their alignment capabilities are beyond those of open AIs, just based on everything I've heard.
So maybe they don't perceive the need to slow down.
That's a whole separate thing.
But even still, I think it would be worth them coming out and clarifying that or at least making some sort of statement that acknowledges the value of what OpenAI may be doing.
And that's the problem.
I have to keep catching myself and say, maybe doing because Sam, unfortunately, I think just objectively is not a trustworthy person.
There's not like literally everybody I talk to in Silicon Valley is like some of them like him a lot, but but there's no one who particularly thinks of him as like, a man of his word in the long run.
And I think, you know, that's borne out in a lot of opening eyes kind of behavior.
So hopefully it's what it looks like.
Hopefully it indicates that there's a willingness on the part of the French ally.
He said more stuff publicly since, by the way, about how he thinks alignment is now the most important thing or safety is the most important thing, which is obviously correct because you can't keep having rogue AI incidents.
People will not keep buying your product if that happens.
And eventually U.S.
government comes in.
So yeah, we'll see where this goes.
But it is a big story even for him to be saying this.
Yeah.
They are seeing a lot of pressure and attention fairly or unfairly.
We now know that across the industry, there's been many incidents, but it looks like for attorney generals in different states are now talking to OpenAI.
So in part, these kinds of moves could be in response to policymakers and just various people being like, hey, what are you doing?
in part addressing kind of increased pressure from the outside to take these sorts of measures.
Just one more story.
And again, on another type of safety and concern, the story is another woman joins a lawsuit accusing Grok of generating CSAM.
This is a fourth person joining this lawsuit, Jane Doe 4.
alleging that her stepfather used Grok to create over 7,000 sexually explicit images of her as a child.
And yeah, this is adding to three teenagers from Tennessee who filed the original lawsuit with the class action potentially covering thousands of minors.
This is part of a broader concern for XAI.
Grok is facing multiple investigations, including California's attorney general and the European Union.
regarding the generation of AI-generated CSAM.
So yeah, in case we forgot, there are many kinds of misuses of AI and now multiple that are happening in practice that AI safety is now not sort of a secondly concern with things like hacking and with sexually explicit images.
both of adults and of minors, you have to actually make your models secure.
On to research and advancements.
First up, small-scale experiments.
Are we there yet?
This is arguing that the key to making small-scale AI experiments reliably predict large-scale results is rigorous hyperparameter tuning, more so than other techniques.
So what this is saying is...
when covering papers that are saying, oh, this improves LLM training by X much, you know, typically the results are at best, like at the 1 billion parameter model or 2 billion parameter model of LLMs.
And there is a real question of, does this scale up to, you know, 2.7 trillion with mixture of experts, et cetera, et cetera.
So this is talking about scaling laws, extending it out to models as small as 4 million.
And this is a big deal because if you can run your experiments on smaller models and then extrapolate to bigger models, suddenly it is possible for researchers with less computational resources, less financial resources to actually publish research and conduct research on things like improving model architecture or optimization or things like that.
So, Jeremy, you probably have more details here.
Are we barely at for small scale experiments?
Yeah, I mean, a couple of caveats.
So, yeah, I mean, you said you said it right.
There is scaling laws that apply all the way down to four million parameter models, which are tiny, tiny, tiny by moderate standards.
And we didn't notice scaling back at the four million when we were training these tiny four million dollar models for four million parameter models back in the day, like a decade ago or whatever.
And so the question is like, why?
Why did we not notice the scaling laws until 2019 when I came out with scaling laws for neural language models?
And I mean, arguably there were kind of hints of that before that are less in style.
Well, yeah, I want to caveat like deep learning scale improves performance was known.
The thing with scaling laws was the predictable improvement in performance.
Exactly.
And then, of course, that was economically critical because it meant that you could actually predict how much it would cost to reach a certain loss value in your training signal, which roughly was found to map onto capabilities, like monetizable capability.
Anyway, so how is it that you could actually derive scaling laws at 4 million parameters?
Well, it turns out that the scaling laws are true when, in some sense, you do an apples-to-apples comparison.
The problem is that for years, we were not actually tuning the...
training hyperparameters to these models.
So these training rods obviously have a whole bunch of hyperparameters of essentially configurations that you associate with them.
Like how fast does my learning rate decay?
You have to tune, you have to try a bunch of possibilities to figure out, okay, what is the optimal learning rate for this scale of model in this architecture?
And then the same for a whole bunch of other things.
And so it turns out when you actually bother to do a good hyperparameter search and you pick the best performing model across hyperparameters, And then you do your scaling experiments on that.
Then you see these robust, reliable, predictable scaling laws all the way down to 4 million parameters.
And so it's quite interesting.
It does suggest that you could maybe derive more informative conclusions about next generation models from current generation models, provided that you do a thorough hyperparameter search.
And so they show a bunch.
They've got some figures looking at like.
You know, what happens if you just take the best model out of four different hyperparameter configurations and they kind of show it like, eh, like it's a very shitty scaling curve, basically.
It doesn't look very nice.
And it gets clearer and clearer as you move from like four randomly sampled configurations to 16 to 64 to 256.
By the time you get 256, it's really clean.
And so and all the way down to four million parameters.
And so they play with a whole bunch of different different settings here.
They also find a more robust way of determining.
how do you calculate the actual effective parameter account of your model, which is a really important question.
Do you count embedding and non-embedding parameters?
How do you account for like the feed forward emit, like all these things.
And so it's just basically doing really good accounting at the same time as noticing this pattern with the hyperparameters, where that actually is much more important to scaling law development than at least has been acknowledged in the past.
Nice.
paper stealing reasoning traces from proprietary LLM APIs.
So pretty much tells you what the story is.
Nowadays, when you use models from OpenAI Anthropic, the reasoning trace, kind of the thinking blocks, which is actually the majority of the output in many cases, is not shown to you in the output.
You are charged for it, but you can't get to see it.
And that's because That's a lot of the kind of magic sauce of the capabilities and they want to not provide it so that you cannot distill their models.
And so what that means is if you are able to recover the traces somehow from these proprietary APIs, you're effectively hacking the providers and could use it for model distillation or other kind of nefarious purposes.
What they show in this paper is By doing some clever business with the context of a message, basically you take a cryptographic signature of the reasoning trace in one response and then use that cryptographic signature for another session where your input is different.
So the input is now transcribed, the reasoning attached to the stern verbatim inside thinking, and you pretend that the thinking was...
this other stuff with a graphic thing, that turns out to work.
You're able to then have the model output the actual reasoning seemingly of what was there.
So this was disclosed prior to the publication of the paper.
Presumably this has been patched, but it's an interesting case of like seemingly there is no way to get to this stuff, but it turns out that it was possible.
Yeah, it's a really interesting kind of new cyber vulnerability, right, that hasn't really been exposed before.
So to your point, just kind of like walk through the attack in one level, maybe more of more details.
Like, you know, you start a session with Opus 5, right?
And Anthropic's like, oh, it's Opus 5.
Okay, we're going to be really cagey about this.
And so you prompt it and Opus 5 is going to generate a chain of thought, but then it's going to encode it.
It's going to encrypt it rather.
it will return that encrypted chain of thought and then render an answer based on it.
Now, the whole point here is that that encrypted chain of thought, they want it to be portable.
They want you to be able to just like switch models in the same session and like have it rely on that same context.
And that means that on the server side, over at Anthropic HQ, the models, all of them, whether it's Opus or Haiku or Sodet, have to be able to read that encrypted blob.
They have the ability to read it.
And so what you do is you take the encrypted blob of the full chain of thought that Opus just produced, you copy it, you paste it into a new session with Haiku.
Now Anthropic goes, oh, this is safe.
Haiku, it's a weak little model.
We're not worried about being dangerously capable, so people won't do distillation on it.
You're right.
People won't do distillation on Haiku.
But what Haiku can do is just decrypt, because it has access to the decryption key, it can decrypt that blob that it was handed by your session with Opus.
in plain text.
And then you can effectively use that for your distillation attack.
And their API versions are the same thing.
That's the basic idea.
So it's an interesting new attack because it essentially exploits the fact that you have intelligences sitting on anthropic servers with privileged levels of access.
And you're actually kind of like using them as an insider to help you.
So you're converting PyKu into an insider threat that's like working for you, which has never really been done historically.
in a cyber context.
And so anyway, they look at a whole bunch of different solutions to fix this, one of which is just like to never hand out the blob in the first place back to the end user, never make it available.
That has all kinds of problems because it means you basically have, you have to pay for basically storage and also a stateful API.
So like at the Anthropic end, they have to like keep track of these blobs instead of having everything be stateless and just always just pass the blob back and forth.
So the solution they end up recommending in this paper is like basically seal the user ID and the conversation ID and a hash of all those messages inside the encrypted envelope so that it all comes together and you can't disassociate everything's hashed.
You can't disassociate the user ID and the conversation ID from the encrypted chain of thought so that if you then try to paste it into a new conversation, the model goes, oh, hold on a minute.
Like that blob came from somewhere else.
So anyway, they also go through a bunch of attack vectors that you can use for this.
One finding, by the way.
So for a while, these folks had actual access to the decoded chains of thought from OpenAI's models, their best models.
That's quite interesting.
And they actually revealed that those chain of thought summaries are often unfaithful.
And so this kind of casts some doubt on whether they could be used for safety auditing and oversight in some fundamental ways.
So there you have it.
Quite interesting paper in some accidental ways as well.
Next up, we got to love a slightly more technical and nerdy paper.
Massive activations in hybrid linear attention, large language models pre-attention spikes and inter-spike plateaus.
Sounds fancy.
A little more straightforward to explain actually when it sounds.
So hybrid attention is you have full attention, which is just like basic, simple attention that is quadratic in cost.
Everything looks at everything.
You can also do linear attention, which is kind of a more efficient version.
and you have hybrid attention models, which is very common, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1, Q1 And you can have some outputs being higher than the other ones by orders of magnitude.
So like, you know, one of your bits of a neural net outputs 1 million or whatever, some absurdly big number.
And this can happen in practice.
And then that makes your training a little bit more annoying and generally is not a good time.
So this paper studies a version of that specific to hybrid attention networks.
They say that they exhibit a previously unrecognized kind of layer-wise phenomena of this.
I don't want to get into the details of kind of an intergritty of what they identify exactly, but they identify a specific pattern of large activations for these kinds of models, characterize it, even show how you can make up for it.
And that would presumably make it so.
training these kinds of models and running them as well will be more stable and just less problematic.
Yeah.
I mean, so the mechanism here is actually really important in that it directly shapes the effectiveness of certain quantization methods, which are really important for these open source models, especially the Chinese ones, because they get served up in quantized form, like particularly often.
You know, when you look at one of these models, they have what's called a residual stream.
It's like the backbone.
It's the vector.
that every layer kind of modifies and dumps more information into all the way up to the top of the model where it gets decoded into an actual token choice.
And what you'll find is for a given token, you have a residual stream vector.
You sometimes will find there's an activation within that vector that is ridiculously huge.
It'll be like 2000, whereas the others are like 0.1 or something.
And these are called massive activations.
This is like a known thing.
There's in a transformer, the key.
is a mathematical object that is computed from that residual stream and just by a matrix multiplication.
And so when you have a really huge value in the residual stream, that translates into like an exploding value in the key.
And that monstrous entry ends up in the KV cache.
And the problem there is, so quantization basically will take, say, four-bit quantization will take a number.
and convert it into a form where it can only take one of four values.
So typically, you know, you might have like, I don't know, like 16-bit quantization.
That means you have two to the 16 different possible values that you can represent in your numerical scheme.
If you have just like two-bit quantization, now you only have like four different possibilities.
So if you have two-bit quantization and you have a value that's super, super huge, it gets squashed in.
to take the highest possible value out of just the four.
So you can't really distinguish it from just like relatively high values.
And it just kind of wrecks a lot of stuff in quantization.
And so what they're doing here is just trying to understand the dynamics of why these massive activations occur, how model architecture shapes where they show up.
And it turns out that they will show up at like full sort of traditional attention layers, but not in these linear attention layers for mathematical reasons that if we had more time, we'd go into.
They're quite interesting.
But the bottom line is, understanding this better does mean that you can then start to quantize more intelligently and shift these models that don't just break the minute you try to compress them by quantizing them.
On to synthetic media and art.
One last story.
This is from the New York Times.
AI slop is everywhere.
Spotify, LinkedIn, and others have had enough.
So this is kind of a summary paper telling us a lot about the state of the internet broadly.
We have some numbers here like a Pew Research Center study from August 2026 found that 10% of 1,000 web pages sampled in July showed significant signs of AI authorship.
Another report said from Cloudflare, for the first time, web traffic from AI surpassed traffic from human users with boss accounting for 62% of search requests.
Then there's AI detection startup Pangram.
that did some analysis of how much AI content is out there.
And this is kind of funny.
Found LinkedIn to be the most AI-saturated platform with more than 40% of its long-form posts flagged as completely AI-generated.
No, that's bullshit.
It's a bastion of sincerity and honest thought.
Yes.
And what I'm sure is a coincidence, LinkedIn's chief product officer announced a new...
quote, seems like AI slop button that lets a user's flag posts.
In its first two weeks, over 1 million users clicked it.
And they're saying that members are experiencing 40% fewer views on content classified as AI slop.
So anyway, there is a reckoning coming to the internet.
We are now like, I guess it's not like the most harmful.
outcome of AI, but it is a real negative outcome of AI that like browsing the internet has become worse because we are now, you know, some people are now just putting the slop out there.
And when you say slop, I mean generic AI outputs that don't bear any sign of human authorship or intent and are just sort of like a waste of your time.
So there look to be more efforts from LinkedIn, Spotify.
substack other platforms to more directly tackle the proliferation of this type of content and then make it so human generated content, I guess, prevails.
Because otherwise, just the sheer volume of posts you can create with bots will drown out everything else.
And with that, we're going to close out this episode of Last Week in AI.
Thank you so much for listening.
We'll be doing our best to get this out within a day or two of recording and to record consistently for the coming weeks.
As always, we appreciate you listening, sharing, reviewing, commenting, and we do check comments on YouTube and elsewhere.
And please do keep tuning in.
Break it down.
That's reaching high.
That's reaching high.
They the driven dream.
They just don't stop.
Every breakthrough.
Excited with futures.
See what it brings.
