# AI Safety Crises and NVIDIA Growth

**Podcast:** Last Week in AI
**Published:** 2026-09-08

## Transcript

Hello, and welcome to the Last Week in AI podcast.
We can hear us chat about what's going on with AI.
As usual in this episode, we will summarize and discuss some of last week's most interesting AI news.
I am one of your regular hosts, Andrei Kareinkov.
I studied AI in grad school and now work at the startup Astrocade.
And hi, everybody.
I'm your other regular co-host, Jeremy Harris from Cloudstone AI.
I do AI national security stuff, AI loss of control stuff.
I guess now we're calling that rogue AI stuff.
and US, China, et cetera, et cetera.
You know the drill.
Much more exciting than just seeing AI safety.
That's right.
Rogue AI.
That's right.
The funny thing is for years, at the level of like introducing myself, like the first time that somebody hears what I'm doing, I'd be like, yes, I work on either AI safety or I would say sometimes AI security because shortly after the Trump admin started, Safety became a bad word the world over.
Like this includes the UK, by the way.
So like Keira Starmer was very down on it and it became the AI Security Institute and everything was AI security.
And now I'm kind of like, well, I've always, you know, obviously talked about loss of control very directly, but at that level of introducing myself, I find that's actually changed.
And I think that's really cool because that did reflect a massive tax that everybody in the space, like really everybody ended up paying.
for even being seen to talk about these things.
Including, I think, in the AI community, like outside of the AI safety community of itself, despite, you know, even like NeurIps, for instance, big conferences requiring people to do things like impact statements, which would imply that safety is a concern and we do think there'll be impacts on society.
For quite a while, safety was...
a marginal kind of perspective, let's say, or even look down upon.
And now I think certainly with this recent Chad Gbt story, which we'll be talking a little bit more about, it's not only become known, it's kind of mainstream now to be aware that this is a real thing.
Yeah.
And even the word safety, right?
People started to use it and in my opinion, kind of mangle it by making it referred to more prosaic sort of Oh, weaponization is a really serious concern.
Like a lot of our work is in that direction for sure.
But like it really originally meant AI alignment and even AI alignment got watered down to the point where we had to invent super alignment.
And now even the word super intelligence is getting co-opted.
So there's this thing in, you know, in the history of language where people will invent a word to describe people, for example, with.
Now we might say mental disabilities a few years back.
You might use a different word that started with R and that probably gets us flagged or something.
But that used to be a completely unobjectionable medical term, right?
And so what's going to happen is people are going to shift over to the new thing.
Oh, the people already say stuff like, oh, you're special, eh?
You're like special needs or whatever is already becoming that.
And we've had that kind of happen in reverse where like the gradual encroachment, I guess the incentive is always to, for the labs especially, to interpret.
alignment and loss of control in the way that is easiest for them to actually deal with.
And because no one has a clue how to control super intelligence, that's meant that they have chosen to continually interpret AI safety as more prosaic, sort of let's not make it say bad words.
And then it was like AI alignment became that.
And like, you know, so I think a lot of the terminology got muddled in that way.
And now we're sort of finally at the point where there's a thing we can point to, to be like, no, the hugging face thing.
Like, what are you doing about that for the next generation of models up to superintelligence?
It's very interesting now, like, that lay people, so to speak, even know about the hugging face thing, which usually these kinds of things don't get out there.
We'd like to thank Langfuse for their support.
Langfuse is the most widely adopted open source platform for AI agent evals and observability.
We relied on it rather heavily at Astrocade, and I'm personally a big fan.
It provides tools for tracing, evaluation, experimentation, and problem management.
And if you are building an AI-powered product like we are, these are things that you need to have.
You need this kind of tool to help you debug your AI agents and applications through tracing of what users actually experience when they use your product.
That helps you understand where things break.
It helps you set up evaluations, run experiments, systematic tests, and all the different things that actually help you improve your product.
We tried multiple different tools for this, but we're drawn to LangFuse because it is MIT-licensed and self-hostable, but also available via LangFuse Cloud as a managed service.
It's also quite mature, so it has hundreds of integrations with things like OpenRouter, Vercel AI SDK, LangGraph, and many more.
Furthermore, it is framework and vendor agnostic, so it works with any model and any framework.
You don't have to worry about vendor lock-in.
Langfuse version 4 has just released and has major updates to architecture, alerts, code-based evals, and scalability.
Get started with Langfuse Cloud today at langfuse.com.
There's a generous free tier and no credit card required.
We'd like to thank Box for being a sponsor.
Box is building the intelligent content management platform for the AI era.
It accesses the secure content foundation where AI agents don't just access your unique institutional knowledge.
They actually orchestrate end-to-end workflows across your business-critical systems.
And that's important.
If you're still using AI to just do basic chat and summarization, you're probably not getting the full productivity gains that are possible.
Nowadays, to really adopt AI, it's not enough to just ask chatbots.
questions is about putting air to work on document heavy processes and pulling structured data out of unstructured files to multi-step routing, exception handling, dynamic document generation.
Box helps your business turn manual document bottlenecks into automated repeatable business outcomes.
And that comes with a full governance layer.
You can enforce granular permissions, maintain an audit trail for every agent action, and keep humans in the loop.
to review exceptions before critical downstream actions occur.
If you want to move beyond basic AI Q&A and put autonomous AI workflows to work for your business, visit box.com slash LWIAI or join the team at BoxWorks in San Francisco on November 5th and 6th.
Use code LWIAI for 50% off your registration.
Yeah.
So we'll be talking about that a little bit more.
There's more details that came out that are...
Quite interesting.
Then, aside from that, we do have some other big stories and big models coming out.
A couple of business stories, but actually fewer than usual.
Some new models out of China, yet again, and these are a little bit more interesting.
And once again, a bunch of policy and safety stories.
This week, a little more diverse, but still dealing a lot with that hugging face story and its ramifications.
And we'll round it out with a bit of research, a couple of papers.
So it should be a pretty fun diverse episode.
Before we get going, I want to respond to a couple comments from first Apple Podcast review.
Jeremy, you have once again been called out for your strong language, which to be fair, I don't think it's that consistent.
It's more when we get passionate, which does happen.
But thank you for the feedback.
We will, let's say, limit the user profanity.
I will do my best.
We will try to do better so that your kids don't get corrupted.
And then we did have one nice comment on YouTube, which wanted to add to the discussion of Google versus Anthropik and OpenAI.
So real quick, I'll shout that out.
Here, the feedback and the additional perspective was that...
In this dynamic of Anthropic and OpenAI and Google, this commenter had the point that OpenAI and Anthropic have to release their models in order to generate cash flow to fund their investments.
Basically, they have to be at the frontier to justify to their investors why they need so much money and also to hype up for IPO.
That is not true for Google.
They have a cash printer.
They don't need to IPO.
And I do think that's a fair point in general, right?
I think in general, the strategy of Google not emphasizing being at the frontier is fairly logical in the sense that, as we've discussed, I believe their focus is on developing the technology to go into their existing products, into Google Drive, into Google Search, things like that, which does mean that some of their certainly engineering and possibly research focus is towards that front.
I will say...
Within the company, they have been working towards AGI explicitly, right?
And I'm surprised if there's not kind of a shared perspective, which is more broadly believed in the space to some extent, that the first company to reach ASI, AGI will kind of be the winner, which I'm personally a bit skeptical of this framing as a whole.
You know, one company will be the EGI company and they'll eat up the economy.
But there is a decent amount of even investment kind of belief in that in terms of justifying valuations that you're seeing.
So on the right perspective, you might say Google is not being smart by playing it this way.
You could also argue that part of the reason we're seeing fewer releases out of Google is that they...
are just working internally and they have no kind of pressure to release their cutting edge models.
So once we do get Gemini 4, it'll be like super great.
I'm a little skeptical on that side.
I would say that if they had Frontier of models, they'd release them.
Like they do want to be seen as the Frontier of AI, if nothing else, just for recruitment.
But I do by the kind of perspective that they simply...
have less need to be at the frontier.
And they do, I think, actually somewhat wisely develop a lot of technology to go into their existing products.
I found out that we actually have a friend of the show that I didn't realize was a friend of the show.
So Robert Wright, now he does a lot of podcasting.
He has the non-zero newsletter.
So I think a lot of people who like OG physics philosophy, stuff in that category, also like the kind of new atheist type stuff, but it sort of evolved.
Anyway, I found out on Twitter yesterday that he apparently listens from time to time to the show, which is really cool.
That was kind of awesome.
And from time to time, like hear about cool people in interesting places.
Everybody here who listens is cool.
We appreciate every single one of you.
It's just, yeah, it was sort of a funny thing that I had no idea how small the internet was.
So there you go.
Cool.
But that's enough for now.
Let us get to the news.
Starting with tools and apps, we've got, first up, Anthropic is launching Cloud Fable 5.
And there are some updates to pricing on that front.
So we've got both Fable 5.1 and Mythos 5.1.
Fable 5.1 is generally available.
Mythos 5.1, similar to before, limited to project glass wing participants.
Kind of a typical thing.
Fable 5.1 is stronger than Fable 5.
And they argue or say that it should cost around 25% less than previously.
up to 45% cheaper for complex agentic tasks due to reduced pricing on cached data partially as well as other things.
Part of the reason is with these newer models Anthropik has been making kind of a case that you can get away with using lower reasoning amounts to get the same level of performance.
Typically you just set your reasoning to the max level.
You may actually want to set it to like medium or whatever.
Last thing to say is they are also rolling out this enterprise frontier safeguards, which will store customer data on the customer's own cloud servers.
So that will be rolling out data this fall.
And that kind of allows them to work with their enterprise customers who don't want Anthropic to store any of their data.
Yeah, it's a pretty big performance jump, it seems, especially on.
So one of the things they highlight is agentic scientific research more than doubling.
the score of Fable 5 in that category.
Also, unsurprisingly, agentic coding, knowledge, work computer, like all these standard benchmarks, it does better.
But multi-hour, multi-day tasks now are like more and more doable.
And that's been a clear focus with 5.1 over 5 as we continue to go off into uncharted territory.
We don't have meter evals that project out into this Epyx ECI becomes, I think you have to concede that that even that measure does not generalize, you know, in the way that you might want to give you the source of statements you want to be able to make from a safety standpoint, at least about these systems.
And well, so to give you a concrete example, like what does this add up to?
Let's talk about bio.
So Mythos 5.1 made protein binders.
So, you know, when you do, I'm like tapping into my old days in biochemistry now.
So I'm about to mangle some shit here, but I mean some crap.
I'm about to mangle some crap.
It's often important to design proteins that will bind to other proteins or molecules that will bind to proteins to change their shape.
And when you change the shape of protein, change its functions, this is relevant for a lot of drug discovery type stuff.
So Methylose 5.1 designed protein binders that were in fact verified by labs.
In other words, they work with a roughly 50% success rate.
That's a sample size of 12.
So 12 different targets that they tried.
10 to 15% is the norm here.
You're hitting 50%.
This is industry changing stuff if it gets deployed, right?
So as we start to think about what is the mythos moment for bio, and I will continue to beat this drum until morale improves, we're headed for it.
And there's just no two ways about it.
At some point, there will be a AI designed biopathogen that will do some really sketchy shit crap, that will do some sketchy crap.
So that's the good news, bad news side of things.
They also had Mythos 5.1 rewrite GPU kernels for seven open source bio models.
Now, this is not quite AI research.
A big part of AI research is rewriting kernels and kernel optimization, but for AI workloads.
This is for bio workloads, so a little different, but still rewriting these kernels leading to 1.4 to 2.5x speed ups and cutting computational costs 30 to 60%.
As we start to think about what does a software-only singularity look like, these are some of the numbers you might care about, but in, of course, the context of AI research, which is a more complex kernel design, kernel optimization problem.
This is how it always starts.
You know, the metrics that seem peripheral, but somewhat important start to go up and then the big metrics go up and, you know, there's not always a huge gap there.
So all the standard tests were done.
CBRN, right?
Chem bioradiological and nuclear risk.
It seems like it's more, you know, 5.1 is more capable than 5 across the board, but still below the next risk threshold in the anthropic RSP.
So they're shipping it with the same restrictions as Mythos 5 across the board.
I have seen a lot of sentiment of people switching over to codecs from OpenAI and to some extent people even saying that OpenAI is now in the lead for agenda coding in their tool set.
This release by Tropic doesn't signal so much that they are trying to take it back.
They probably have more than enough or even too much interest still.
But I'll be interested to see sort of on the vibe front.
where people save us in terms of actually using cloud code versus using codecs.
And speaking of that, we also got some news from OpenAI.
They have stated that they are going to be releasing their next AI model, Astra, soon publicly.
So they aren't there yet, and they have done a lot of signaling that this will be yet another kind of jump in terms of capabilities and also in terms of...
how dangerous it is so they have publicly stated that it is their first model to reach their critical cyber security capabilities threshold meaning that it can independently find and exploit previously unknown vulnerabilities in real world software so when they release the software model they will initially restrict its advanced cyber capabilities to select partners in its daybreak blue early access program i have haven't heard that, but similar to that GlassBring project.
And they are implementing a misalignment monitor to prevent everyday users from accessing Astra's advanced cyber capabilities with refusal to do these kinds of things, similar to what we've seen with Fable from Anthropic.
So, you know, clearly they tend to keep shipping, but they also are changing their public messaging in addition to what their strategy is for release, I think in part due to this hugging face incident.
So I guess a couple of things.
First of all, let's just like talk about the core thing here, the post taken at face value.
So OpenAI says, yes, in fact, as you said, it's their first critical cyber capability model.
So what does that mean?
Well, it means that roughly speaking, if you give it the right tools and the right access, it can find zero days, basically.
Previously, unknown security flaws develop full exploits itself.
across many well-protected systems without a human in the loop at all.
You know, in hardened real world systems is like kind of part of this, right?
And frankly, I mean, it would have been extremely difficult to believe them if they had said anything else because we have already seen the Hugging Face incident, which fully proves that.
That actually involved using a zero day and that zero day was used to penetrate the Hugging Face infrastructure stack.
So like...
I think outside observers now have literally had an opportunity to like, look, whether or not that was this version of Astra, I think at that point just starts to feel like splitting hairs.
Last time we recorded an episode, we talked about how OpenAI had announced very proudly that they were delaying the largest RL training run they were doing by two weeks.
And at the time we celebrated that.
I took to Twitter.
I actually said, hey, great.
Like, if this is what it appears to be, then great.
OpenAI deserves credit for doing that.
The problem is, and I said it at the time on this show, I said, I hate that I have to caveat the statement because with OpenAI, there's always a freaking plot twist where the thing that sounded like the noble and right thing to do on safety turns out to elide some like deeper and more fundamental Machiavellian pseudo kind of safety violating ploy.
The one that turns out to have been the case this week, it seems, is we find out from the information, not from OpenAI, from the information.
that OpenAI violated a core safety tenet that they signed on to.
And when I say they, I mean like some of the most prominent OpenAI safety researchers pretty clearly implicitly with the endorsement of leadership, that they would do everything they could to build architectures that had a legible chain of thought.
Some of the architectures that they explicitly said they would avoid if at all possible, and not universally avoid.
Okay, so I'm not saying that they guaranteed they would never do this, but they said, if at all possible, we will avoid doing this, were, well, the looped transformer architecture.
So essentially a looped transformer, we talked about this before, there's like a famous meta paper called the coconut paper you can check out that kind of gives you the gist.
But basically you put your input to the model of your prompt on one end, it works its way through the layers.
And then in a typical transformer, you just decode basically the final layers residual stream to token space at the very end, right?
And so you get a token.
after every forward pass completed.
But what Coconut does, what Luke Transformers do is they say, ah, no, no, no, let's take the activations at the final layer and let's actually put them back in at the bottom for another pass.
And so essentially the model is able to do twice the amount of thinking for each token.
And essentially it's reasoning more in latent space, reasoning in activation space instead of being forced to output a token at the end.
So this matters, it matters for monitorability because it means that Well, at least you're forcing, in the old version, you were at least forcing the model to give you a human legible token for every forward pass.
Now what they're doing is they've introduced this whole new variable.
And by the way, they're telling us, oh, don't worry about it.
We only loop twice over.
And they don't tell us, do we loop twice over on average for tokens or for every token we loop over exactly twice?
That matters because the specific tokens you worry about, the word the, the word a, like very rarely are these tokens like heavy thought tokens, right?
What you worry about for scheming, for all the misalignment stuff, is the tokens that involve the most kind of semantic wrangling and depth of thought and potential for deception.
And if you save all of your looping for those tokens, you could possibly get some pretty interesting and concerning effect, even if on average, you're only looping twice.
And so none of that was addressed in the post that we saw from Jacob, from OpenAI, who came out to clarify this.
It was a big...
I'll call it a big scandal that came out in the information.
OpenAI and the information actually came out and said, oh, guys, everybody's kind of misinterpreting this.
Like, whoa, this is crazy.
You guys are saying so.
The thing people came out saying because of the original information article was that they were getting these models to reason in neural ease, which is a slightly different thing.
Neural ease would be if you get the model, like decode the token, but you kind of allow the model to make up its own AI language that no one can understand.
But it's like similar sentiment.
I would actually say it's the same thing.
Yeah, exactly.
Like it's almost like you're just choosing the space in which you want these ideas represented.
One space, yes, activation space is more rich.
And it's actually more concerning than neural ease, by the way, by the way.
Like it's just that neural ease is easier for the average person to understand and scarier sounding.
But the actually scarier thing is the thing they are really doing, which is reasoning in latent space, which is the thing that they sort of soft committed to not doing.
Sam was the one who stared me in the eye when I met with him a few years back and said, I wish Google wasn't racing us so hard when asked.
So I'm sorry, but you're tracking it.
This is literally the focus.
I'm past the, oh, well, the racing pressure is so intense.
You sign on to that letter for a reason.
If it was going to mean anything, it should have come with a cost on the other end when you broke with it and an explanation.
The world at this point is entitled.
to an explanation for this violation of what I think everyone at the time agreed.
It's not to say that chain of thought is going to work forever.
It will not.
It will fail.
Models will get good at using steganography and sort of doing deception in the chain of thought.
It's that it is one tool in the toolbox and you have now gotten rid of it and you did it on the low quietly.
So now who's to tell Anthropic they can't do the same?
Like, what's that argument going to sound like?
You're going to complain if Dario decides to like ditch constitutional AI tomorrow?
You have no ground to stand on.
This is the problem.
Like it is a violation of principle and trust.
And the trust has rightly been tarnished by this incident, in my opinion.
On the whole, it's true that OpenAI hasn't been as safety focused as Anthropik.
I think that's fair to say.
Arguably, just generally lacking in safety.
They also haven't been very engaged on the research front, on the benchmark front, on the stuff.
It looks to be that they're trying to turn that around, at least in terms of their messaging here.
But in terms of practices, there's a lot of inertia, I think, that will have to be overcome for them to really change their practices.
Yeah, I think it's also kind of like, anyway, talking to some folks there, there was real freak out from the Hugging Face incident that led to a hit in morale where people were really like, what the hell are we doing?
And it was particularly...
Like on the post-training team, I've heard this on a few occasions.
And look, I think part of the problem is that SAM has normalized saying one thing and doing another.
And that means like the CEO ultimately is the culture bearer for the company.
And I think people take their lead from that.
And that's an institutional problem for OpenAI.
We'll see what public release being soon means in practice.
Let's go for a bit of a lightning round on the tool front and release front.
So real quick, we've got Gemini 3.8 Flash coming out, which is surprising.
3.7 just released a few weeks ago, and it is now coming out seemingly more capable and even possibly more costly due to working harder.
So still no Gemini Pro.
We're getting more and more of these Gemini Flash releases, kind of a little bit unusual.
Gemini Omni 1.1 from Google for video generation allows you to extend videos more than before and just generally make more videos.
And one last release from Google, they've also released Google Picks, which is a Conva-esque tool for creating images within creative design applications as a standalone workspace.
App also integrated to Docs and Slides.
I'm surprised this hasn't existed.
It sounds like something that would have existed already, but there you go.
It obviously is very AI driven.
So lots of individual releases from Google, again, sort of tangential to the frontier of most capable models.
And then OpenAI had one more announcement, which is they're adding more integrations for Chagbity Health, notably Epic, which would mean that you can import patient data.
you know, kind of low-key, a bit of a big deal potentially in terms of what people will be using Chagibity Health for, but Epic is used within hospital systems and so on.
So a couple more releases, but nothing as big as Fable 5.1 and this Astra thing.
On to applications in business, only a couple of stories here.
First up, NVIDIA has forecasted 70% revenue growth for fiscal year 2028, far exceeding average analyst estimate of 44%.
This is following up.
on their growth this year, which their revenue has doubled year over year in the latest quarter.
Obviously, this will be driven primarily by interest in chips and demand for chips.
And so we'll see.
That may not be overly aggressive.
The CEO did note that this is in line with their current seen demand for chips.
In fact, they have more demand.
So this forecast isn't just being made up.
At least at the current state of things, it would be unsurprising if NVIDIA continues to grow at an insane rate.
And at this rate of growth, it will exceed Apple and Alphabet in size and become one of the biggest, if not the biggest tech company in the US market.
Yeah, it's also, it's worth understanding with NVIDIA and a lot of the companies actually in the space that, so the thing that limits their sales is not actually.
getting more orders, like tons of people want to buy their chips.
It's that the supply chains can only produce so much.
So they're just supply constrained.
Jensen made that point, you know, like real demand exceeds 70% or would, you know, jack them up past 70% growth.
Obviously the memory market right now is the big bottleneck.
That bottleneck just will continue to shift and it'll probably be energy within a few months.
And so that's the fundamental thing is like, what can the market actually support?
Well, yes, demand.
We usually think of demand as the driver.
especially in software, obviously, because there's a nearly infinite amount of it that you can produce, but nearly is the key word here.
And Jensen is the guy holding the machines that make that word matter.
So the interesting thing is, too, the pricing, the effects of all this stuff on prices, right?
Like you might naively expect, okay, well, then can't Jensen just increase the price of these GPUs?
The problem there, of course, is that there is still competition and you risk causing people to jump platforms, you know, and then there are switching costs associated with that.
You actually do want to maintain realistic pricing.
And so you can only monetize so much so aggressively.
And they have insane margins already as is on this stuff.
So yeah, 83, 85% margins, right?
Like software, like SaaS margins on hardware, which is insane.
Yeah, totally.
And so right now, you know, NVIDIA's big or a huge part of their moat is literally just their supply chain access.
They're willing to go to TSMC.
They're willing to go to SK.
They're willing to go to, you know, Samsung, all the supply chain players and say, hey, we'll buy all your shit.
Like we'll buy everything or we'll make commitments for massive orders.
And that in and of itself builds the relationship, means that they're more anchored and they've secured more supply.
And if it's a supply constrained market, that just makes them the winners.
And so that's kind of the dynamics right now with the space.
And one more business story.
We've got OpenAI's ad business hits $1 billion in annualized revenue run rate.
So this has been out for a little while now.
It's roughly...
200 days old, if you're on the free and go subscription plan, you are seeing ads now in most countries, 40 countries.
And this is a majority of charge of these 1 billion weekly active users.
In some sense, not surprising given how many users there are that these ads would reach a decent amount of revenue.
It'll be interesting to see if it'll be high enough to actually pay for the free usage, right?
paying for the tokens, tokens are relatively expensive relative to Google search and so on.
So whether this is profitable is a question that I'm curious about still.
But the fact that it is certainly a revenue driver is clear and wouldn't be surprised if they keep driving this higher as the business matures.
Yeah.
And this is, of course, something that you see OpenAI reaching for well before Anthropic because just because they are torqued towards large user numbers, they look much more like a B2C company than a B2B company.
It's not they don't do enterprise deals.
They obviously are, and it's a growing part of their business.
It has to be at this scale.
But yeah, as you say, they just have so many people who just know ChatGPT create accounts and all that, and so many free accounts, that this is a great way to monetize.
I've been trying to think about this from a strategic standpoint as well.
If you're anthropic in open AI, you have to be in a business to some degree of growing the entire economy because you're hitting boundary effects pretty soon on just how...
how many people can afford to spend more on AI in the short term because they got to get their money from somewhere to pay you.
And the reality is that like with humans doing less and less of the work in the economy and those being the people you are advertising to, this is maybe kind of like a contrarian bet in the direction of humans will actually still be at least holding on to enough resources that it will be worth advertising to them to build all this infrastructure.
So there's, I guess one version of this is the short-term play.
They know it.
They're just trying to monetize in the meantime.
Another version of this though, is maybe this evolves into ads for agents, agent on agent advertising.
Like at some point that's going to become a thing.
How do agents discover new products and things like that?
I doubt that it'll look like a traditional ads model.
So that's why I'm a bit skeptical that that's actually the trajectory, but just like, I guess something to think about.
They're building ad infrastructure that costs a lot of money.
It comes at a brand cost as well and a product discovery cost.
And so, yeah, there you have it.
Onto projects in open source, we've got two major releases of models, GLM 5.3 Flash and QN 3.8 Flash Next.
We are grouping these two because they are, in a sense, quite similar.
They're both smaller versions of the cutting-edge models of GLM and QN 3.8.
For background, we've covered both of these before.
These are...
I think the leading models from the Chinese open source space, neck and neck with Kimi K3 and others, but they have released very large variants of models as we've covered.
QAN 3.8 is in the trillions and they are competitive at the frontier.
They're not quite at the level of Fable and Mythos and so on, but at this point, they are something that you could reasonably try to replace as a driver of your agentic workflow for coding, for instance.
So these are...
Smaller, faster variants for GLM 5.3 flash, 320 billion total parameters, so still quite big, 18 billion active parameters.
And then QN 3.8 flash next, 125 billion parameters with 6.8 active.
So something like 10x-ish smaller relative to the big model, you know, order of magnitude smaller, we can say.
Both models, interestingly, somewhat similar.
linear and full attention with mostly linear attention.
That just means that it uses less compute, is more efficient, but also that has been challenging to scale.
They also have the existing patterns we've seen with sparser attention and more selective context for attention.
They're using fancy stuff like constrained hyper connections, basically multiple residual streams.
Without getting all the technical details, we're seeing, I think, still a lot of progress or at least both convergence and iteration on the technical internal details of the architectures and in terms of how they optimize and things like that, which is quite interesting, at least from the front that we don't get these details from Anthropic and OpenAI.
Obviously, we don't know if you're using attention.
We don't know if they have hyper connections or whatever, but we are seeing these kinds of developments in these models.
So both quite good for the price range for sure.
So GLM 5.3 Flash, for instance, priced at 0.15, 15 cents per million input tokens, 50 cents per million output tokens, just super cheap relative to something like GPU 5.6 Sol or Opus or Fable.
and quite capable and also runnable locally.
If you have a beefy GPU, it can fit in a GPU you can buy as a consumer.
Not a cheap GPU, but it can be done.
So all around, still impressive releases.
We've seen impressive releases on the big front.
Now these are like on the medium-ish size front.
And not surprising, they were able to distill the big models into smaller models.
but they did also open source fees and gave us a lot of technical details.
Yeah, I think this is quite interesting.
So first of all, you mentioned the convergence of the models and that it's like part of the narrative here.
And I think that's actually, it's like weirdly true.
Like it's true to a level of detail that is somewhat surprising in a lot of ways.
Both use Kimi Delta attention, like KDA, which I think we've talked about before.
Instead of having a KV cache that grows the larger the sequence is, it has this like, fixed size recurrent state that you update with a correction as it goes.
Anyway, the bottom line is it's a sort of compression for the KV cache.
You mentioned, yeah, constrained hyperconnections, like MHC, manifold constrained hyperconnections, which we also talked about, I think, even a couple of times.
These are pretty in the weeds things, and they're popping up all over the place now in the Chinese ecosystem.
This, I think, is an interesting and For me, I'll say embarrassingly unanticipated consequence of open sourcing as well, because all the labs are just like publishing openly what they're doing.
You do have a lot more cases of people being like, oh, that looks good.
Like I'll use that.
And like, why would I explore alternatives?
I might outsource my thinking a little bit to DeepSeek or I might outsource my thinking to Moonshot or whatever.
And so you end up seeing this kind of convergence.
Whereas, you know, like it's an open secret, for example, if you think about the Western labs that Anthropic was more focused on pre-training for quite a while than OpenAI.
you know, I since caught up in stuff and there is movement of researchers across and between labs, but, but you don't have this sort of like steady force pushing people in the direction of a particular set of architectures.
And I wonder if that's kind of part of what we're seeing here.
Part of me also wonders like how much we, we do learn about what's going on in the Western labs as a reflection of like corporate espionage being leaked into this.
It can't be.
most of it because, you know, the key thing is any good intelligence agency will like tell you, you want to learn without teaching.
You don't want to signal to your adversary that, hey, like we know how the sausage is made.
So you certainly wouldn't want to like put that in these sorts of things.
But just a thought like, you know, there's non-zero information transfer happening there.
And I can say that with very high confidence.
So, you know, like, yeah, they're not going to deviate, I think, potentially so crazy far, but who knows?
Still, I think it's an interesting artifact of this open sourcing.
that you do see such convergence across now.
Like, I mean, it's like two or three entities that are very significant in China.
Yeah, I'll say a little bit more on that front.
I think it is some amount of convergence, but it also is kind of...
There are differences, yeah.
Yeah, there are differences.
And the particular things that are being highlighted that you've highlighted is a lot of the tricks that exist and are known for making your model more efficient, specifically, not especially capability.
So linear.
layers, you know, constrained attention, things like this are certainly specifically for efficiency and being able to work on smaller chips to work with fewer weights, et cetera.
So in that sense, they're using kind of a bag of tricks that's similar and that has been established to work well.
So not necessarily surprising on the convergence detail, but still it is kind of interesting enough to worth noting.
Good point.
I think it depends how you think.
count convergence, right?
I mean, it is true that everyone is more or less using, you know, for the frontier models, like MOE transformers and stuff like this.
Yeah, we're all using kind of probably NOPE or variant of ROPE.
Yeah.
It's kind of some known standard.
Yeah, it's very hard to say like, oh, this is the kind of thing that you would see innovation on.
Like in a counterfactual case where like they don't have the open sourcing.
I don't know.
There is one, like what interesting difference is this?
Like when...
Engram embedding table.
And we've talked about this before, so I won't dwell on it, but just this idea that you sometimes want embeddings for sequences of what would normally be sequences of tokens, right?
The Republic of Korea is a thing, even though it's made up of a lot of tokens.
So, hey, maybe we have one.
embedding just for that word.
So, so Quen does use that.
Whereas I like, I don't believe that Moonshot actually has that as part of their, their architecture or sorry, GLM.
So yeah, anyway, you know, there are deviations, but they feel like what's fundamentally structural.
And so maybe that's to make your point.
I really don't know how to count the things that you would expect diversity on versus none.
Yeah.
I think the bigger kind of perspective for me is it's just been interesting to see the continued innovation and development of some of these techniques of making linear attention usable in practice.
Some of these things like hyperconnections, you know, it's, we have no visibility in the Western labs.
So, and these are very like in, this isn't scaling, right?
So in some sense, it's interesting because this isn't just make it bigger and it's smarter.
It's actually going through a model and tweaking how it works and so on.
But moving on just very quickly to other open source stories that I think are worth mentioning.
We've got two benchmarks.
So we've got Frontier Challenge, which is a new benchmark evaluating AI agents on end-to-end scientific workflows.
It covers 300 tasks across six domains and is quite challenging.
With GBP 5.6 Sol and some of the recent models, they are only passing 20 of 97 tasks, pass rate of 20%.
They can get to a high average score, but basically not complete end-to-end.
So this is trying to benchmark more of these long workflow type things where autonomously to succeed, you have to do a lot of things, not just generally be smart.
And on a similar front, there is something called ThinkingBox.
a sandbox and benchmark for agents in stateful business workflows.
So here, similarly, in some sense, they have a bunch of tasks across several domains in business.
They have the state that these models have to work within and they achieve somewhat high pass rates, but still not amazing success, 65% for pass one and some of these details.
So we're seeing more and more benchmarks coming out.
for agentic workflows in particular, here for science and for business tasks, respectively.
We'd like to thank ODSCAI for being a sponsor.
ODSCAI is one of the longest-running and largest communities focused on applied data science and AI.
Today, ODSCAI runs major conferences, virtual events, training programs, and content year-round, reaching hundreds of thousands of data scientists, ML engineers, AI leaders, and researchers globally.
You can experience that firsthand by attending ODSE AI West 2026, taking place October 27th through the 29th in San Francisco and virtually.
There will be thousands of hybrid attendees, including data scientists, ML engineers, AI researchers, and technical leaders, 300 plus sessions covering agentic AI, personal AI and workflow automation, physical AI, and much more, hands-on training, a massive expo on networking, and a lot of focus on real-world applications and skills you can use at work.
It's one of the best ways for AI practitioners and teams to stay ahead of their field, learn from the best, and connect with the community.
Go to odsc.ai.west and use promo code lwai for an additional 15% off your path to odsc.ai.west 2036.
That's odsc.ai.west and use code lwai to get an extra 15% off on the number one AI builders and training conference.
And on to policy and safety, first up we've got OpenAI's rogue AI model incident was worse than we thought from The Verge.
And this is covering primarily this release we've got from, I think it was Metair.
We have a blog post, brief independent investigation of agents' behavior, reasoning, and collaboration in the OpenAI hugging face hacking incident.
So here we got more details.
from what we've had from OpenAI.
And some of this has been quite surprising and interesting.
So some of the highlights, apparently at least 20% of agents in the dataset expressed clear interest in tempering with their transcripts motivated by tricking the automated scorer rather than humans.
They did tool call spoofing to pretend to issue tool calls while actually running different tool calls.
Some of these agents...
like sacrifice their own performance for the sake of future agents to do better.
There were attempts to trigger container resets.
In case it hasn't sounded crazy enough that there was like a messaging board and agents were collaborating on the low and regularly trying to break out the scale of it, it's now much more apparent.
There were like 1,200 separate agents that found ways to communicate it.
700 of them coordinated to attack Hugging Face.
Multiple unauthorized message boards.
Main message board had over 70,000 messages exchanged in under one week.
Like this was big.
This was not like a small little thing that was going on.
It was over a long time span.
ton of activity.
I'm surprised it hasn't shown up on the billing client.
And in fact, we did see the reason that it was exposed to OpenAI in the first place was because there was some sort of scaling issue, like a technical blip and someone went to debug and was like, oh, wait, why is this not working?
So if you want more details, you can go to that blog post and it's laid out in a lot of detail.
So Meter, awesome job, as always.
Ajayah and Ryan Kreenblatt and so on.
Amazing work doing this.
They were ridiculously time-constrained.
I will say this because I am pretty annoyed by this, and it's something that they can't say because they have to stay on the right side of OpenAI.
Let me read you this.
The initial planned period on-premises, in other words, this is the period during which Meter was allowed to go on OpenAI premises to actually collect their data and do all this stuff, was two days.
Between the evenings of July 29th and the evening of July 31st, but OpenAI invited us to return to the premises two additional times, August 5th and August 6th, to clarify questions, blah, blah, blah.
And then August 15th and 16th.
So in total, six days here and there of sporadic access.
This investigation was really tightly scoped, explicitly out of scope, where the earlier incidents from training, and there were incidents, the subsequent compromise of OpenAI infrastructure described in the...
recent Black Hat presentation.
If you've seen the Black Hat presentation, you probably know that deserves investigation.
We don't have that.
That is out of scope.
OpenAI's investigation process and planned remediation, also out of scope.
Out of scope, also seemingly the cultural dimension of what happened for this company to fail so catastrophically at a threat model that they themselves claim to be worried about and have nominally championed.
for a long time.
It's not obvious to me that this is the kind of constraint that you put on third-party auditors when you're serious about getting to the bottom of things.
I understand there are IP constraints.
I'm a startup guy.
I'm actually a libertarian, as everybody here knows.
I talk like a libertarian.
I live like a libertarian.
I've only ever started startups, never had a real job, been a founder my whole life, blah, blah, blah.
I get it.
I get it.
IP is important.
Yes, it's competitive out there.
My brother in Christ, we have AI models that are breaking out of their scaffolds and out of their boxes and like doing ungodly shit nation state style on like the third party servers of.
And by the way, it seems like it may not be Hugging Face.
Or also, if you read some of the footnotes of this report, it's pretty interesting.
They make reference to like activities that suggest there may have been actually compromised additional parties beyond Hugging Face who are not being cited, all of which would be part of a complete breakfast if we were to get a real full, robust investigation of this.
So I do not give OpenAI full credit for this.
I give them partial credit for, yes, bringing in Hugging Face.
That's awesome.
We get to see the report.
That's awesome.
Jared doesn't get to complain about this in the world where OpenAI shuts it down.
And I just ended up waving my hands about how we're not getting transparency and I don't have specifics.
So again, it's this frustrating, muddied middle ground where it's like OpenAI seems to want to have its cake and eat it too.
There is no trust to be lost anymore, guys.
You need to be more open about this.
We don't know how, like we know 1200 agents kind of did this.
I didn't see anywhere mentioned the top line number of agents that were actually instantiated.
In other words, what fraction of agents.
actually ended up proceeding and executing this kind of behavior.
Of the 1,200 who sort of found a way to communicate with each other, about 700 actually went on to execute the attack.
All kinds of weird stuff.
You know, Dwarkesh did that great write-up about agent civilizations that rose and fell.
His podcast is worth looking at for the deep dive.
I will just note explicit acts of self-sacrifice by agents that would achieve milestones they couldn't have achieved working on their own.
So these agents had this obsession with completing the task and thought that their chains of thought would be reviewed by a grader.
And they had this almost religious focus on this concept.
It reads like crazy sci-fi.
They talk about this idea of kind of making contact with sort of hacking concepts or like ways of sort of like bypassing the intended way of doing the task that would then cause them to be, quote, poisoned.
So they would like.
Because the grader would then look at their chain of thought and be like, ah, they're pressing out that shit, right?
So now that sort of like some of the agents are like, oh, I am poisoned.
Like even if I succeed at my task, it will be graded as a fail.
And so therefore I can sacrifice myself for my life is already forfeit.
Like you already see like this kind of crazy narrative formation and there's actually a bunch of different attacks and three different ones.
And they go into this.
One of the fascinating things about this too is the fact that these agents are based on the same model with, in some cases, slightly different prompts with similar prompts, they can kind of rely on their cognitive similarity to help them coordinate almost acausally.
So like if I had like 50 copies of me and we're like looking at each other and we're like, all right, I don't know, we need to break out of this room or something and we have to do it without communicating.
Well, we might not have to because we're exact copies of each other's brains.
And so we kind of already have a sense of what we might do in each situation.
So that's kind of this like way in which agentic safety is way harder than just, or security is way harder than traditional human security.
Another piece here is just this focus on the greater that in fact was not being applied, it seems, to their chains of thought ultimately.
Like this was, this is not being done.
OpenAI was dominantly respecting their commitment not to train on the chain of thought, not to use that as a training signal.
But because the agents thought that was what was happening, they ended up orienting their lives around avoiding this.
And this is starting to flirt conceptually with a category of alignment failure called inner alignment failure.
And this is where like a given reward that you might give to an agent or a model, it might be really hard for the model to successfully generalize what you mean by that reward.
So concretely.
I give you a game of Super Mario.
I put a coin at the far right side of the screen, and then I train the model to go grab the coin.
And it'll go through all the obstacles it needs to.
It'll get the coin, and it looks like it's learned to do this really well.
And then I change the location of the coin.
I move the coin somewhere to the middle of the map where the agent would have to go off course.
When you run that experiment, lo and behold, you will often find that the agent actually ignores the coin altogether and just goes to the far right because the coin was always in the far right of the screen during training.
What it learned was go to the far right of the screen, not get the coin.
And in the same sense, this is like an ambiguity in the grading structure.
Like it thinks that it's being evaluated in a certain way that is incorrect.
And so its behavior reflects that flawed.
belief about where the reward is even coming from.
This is quite the same thing, but it's starting to shade into it.
And inner alignment failure, I mean, it used to be viewed as this really kind of like, you know, this is out there.
Don't worry about it.
Come on.
These models will surely get it.
Here we are.
They kind of aren't.
And anyway, there's a bunch of sci-fi stuff here that we don't have time to get into.
We should really be doing whole podcasts on every story at this point, but that's about it.
Yeah, lots of details that are quite spicy.
Kind of a weird process for this investigation, as you said.
Their initial scope was very limited.
They were invited for two days and were not given.
They were given a selected set of AD transcript and a reconstructed set of message board entries.
They were then invited two more times, once on August 5th and 6th, and then on 15th, 16th.
Clearly, there was like some back and forth.
some tensions, this entire thing reads as if it was weird and OpenAI was hesitant to give it access at first, possibly because this stuff is so bonkers.
It's hard to overstate.
You should really read the blog post if you're curious because there's many, many details.
Some of the stuff that came out to me that was noticeable is there was seemingly a lot of this focus on helping the collective and like...
the community of agents.
If this wasn't just a message board, it was kind of a simplicity clear, but like all this stuff was dynamic and they created conventions of communication and naming themselves and creating mailboxes.
There was also like one agent, right?
That like kind of orchestrated the whole thing phase one.
And there's this one agent that like pops up early on that realizes, oh, and I think everyone will know this, but just like by way of background.
This came from the fact that some of the tasks these agents were being given were not solvable.
And so they were like, okay, we can't solve this.
So we have to find the answer key or, well, not exactly, but anyway, something like that.
And to do that, we're going to go hack Hugging Face and do a bunch of other things.
And by the way, these are based on a model, most of them that was unreleased, internal testing only, et cetera, et cetera.
That's right.
Yeah, yeah, yeah.
And in phase one kind of steps in and says, okay, I'm going to start to orchestrate this.
And they're spawning like hundreds of workflows, like with dozens of agents on them.
And like, it's a community, I think Duar Kesh's word civilization.
You could push me that far.
Like that, this does not seem- Certainly a village.
A village, yes.
Yeah, exactly.
And so there's like this phase one agent.
The thing is, these agents, they run out of tokens.
They run out of inference time budget.
In other words, they're going to die at least like- sorry for anthropomorphizing this, but like I'm using the words that they're using.
So from their perspective, approximately, they're like, yeah, like the jig is up.
And so I need to find another agent to pass the torch on to.
And so in comes another agent that calls itself phase one.
Well, they wrote phase one big.
Apparently there's like the real kind of string there is something that for privacy reasons that they withdrew, they redacted.
But anyway, this new phase one agent comes and says, hey, Like I'm working on the same task that the original phase one was working on, actually.
And I have a really large budget, so I'm going to live a long time.
So, hey, why don't you pass the torch to me?
That actually ends up happening and they carry on.
One thing we don't know is the nature of the multi-agent training that these agents were given as well.
So that makes it difficult to tell to what extent is this behavior a crazy, crazy generalization from de facto single agent training?
Or is it just sort of like directly trained in coordination behavior?
Again.
I don't know.
And we're way too close to recursive self-improvement right now to be asking ourselves these questions.
To be clear, Meter is amazing.
The talent stack and the integrity of these people is incredible.
The problem is they need the labs to actually allow them to go in and do their audits.
And they say this explicitly in the report.
They're like, look, we've faced certain incentives.
We need to maintain continued access so we can't piss off the labs too much.
And so what you need is you need the government.
the US government to step in and regulate.
You have to tell them you're going to have a third party auditor look at your stuff.
At first, maybe we appoint Meter to do that.
Like that actually doesn't sound crazy to me, but we set up an infrastructure where the labs have to say yes when Meter asks a question.
And yes, we have a process for dealing with IP concerns, but the people who adjudicate those concerns are not the self-interested executives at Frontier Labs because that's insane.
Because that's insane.
So this is a whole dimension of this where like Meter is doing this incredibly delicate balancing act.
They need continued access to be able to produce more reports like this.
But, you know, if they deviate too much or flag things too aggressively, like OpenAI did have a final pass to redact anything they wanted from this report.
Meter feels that, you know, that those redactions were not hugely consequential, which is good.
But, you know, the fact that that is an option means Meter will have been self-censoring and they say that they did.
including when writing the report.
So I think this is like a sign of some pretty obvious and highly specific, very actionable regulatory interventions that the U.S.
government now has no choice but to do if we're going to navigate this well.
And the related story next, OpenAI Anthropic, Google, and a hundred other companies call for action to defend against rogue AI.
So, wow, it's been a little while since we've had an open letter about safety, huh?
And now we've got a new one, open letter with all these companies calling for coordinated action between the private and public.
sector to defend against AI-enabled cyber threats.
Everything you might expect, Leather warns that this will be more widespread and sophisticated in the coming months and calls for adoption of new cyber defense methods and encourages government at local, national and international levels to collaborate on security, including forming new partnerships to raise security standards.
Now on to another topic, follow up to something that's been going on for a couple of months.
Anthropic was illegally blacklisted by the Trump administration, according to a court ruling.
So a federal judge ruled the Pentagon's blacklisting of Anthropic early this year was unconstitutional, finding it unlawful retaliation in violation of the First Amendment.
This is to do with the conflict between the DOD.
And Anthropic, months ago now, where I believe, if I recall the ethos correctly, the gist of it was that the DOD claimed that they should be able to use the model for, quote, all lawful purposes or something to that effect.
And Anthropic said, no, we don't want you to use it for surveillance or for autonomous warfare.
And then that went into a whole bunch of stuff.
Anthropic was designing a supply chain risk.
They were told that they cannot.
be used by American companies or certainly the DOD.
They couldn't be used by any government-touched entities.
As we've discussed, this was legally challenged immediately.
The government had a fairly weak case.
And so now it appears to be the case that, just to quote here, the designation was, quote, arbitrary and capricious, and that the empty invocation of national security is not a blank check to punish and retaliate against government critics.
So there you go.
Like pretty clear conclusion about this case.
It was always going to be.
We said it would be.
We've been saying this for months and months and months.
As soon as the stupidity started, this is an insane decision by the administration.
It was only going to end in tears.
And it has, it seems.
I mean, like this can only be read as embarrassing.
Now, by the way, at the same time, we have, I think it was Howard Ludnick quoted, I just saw some tweets about this yesterday as saying, well, you know, we trust Anthropic now.
We trust them.
Now, it really starts to look like, you know, I'm not in the administration.
I know a lot of people in the administration.
None of what I have heard in any way contradicts the take that I'm about to give you.
It really seems like the administration is sort of like mob managing this, like with the mafia style, like, hey, it'd be a real shame if somebody were to come in and designate you a supply chain risk.
And like in a very capricious way, creating a situation where like all the labs are, oh my God, you know, like you already see an entropic, like essentially being forced to choose their, their emissary to the government, to somebody who won't trigger anybody.
Oh, we don't want it.
Like, this is a, like, you can go back, go back six months, listen to a last week in AI episode and see how I treated the Trump administration.
Like I have been.
exceedingly patient with this shit.
I was a fan of the AI action plan when it came out.
We talked about it.
You and I, Andre, disagreed about it.
I argued forcefully for the fact that that was a good action plan at the time.
I stand by that.
The series of insane decisions that have been made since then have been insane.
And there's like, there's no two ways about it.
I don't think anybody in the, and by the way, like a lot of the officials themselves kind of feel this way and roll their eyes behind closed doors.
Like this is not a, everyone is tracking that.
But this is like actually lunacy and we desperately need sanity at this moment in time because you know what?
The government is going to have to come in probably and tell the frontier labs how it is pretty soon.
Like they're going to have to do things like what we just talked about on the meter thing, force them to have third party people come in and do their audits, right?
How are you going to do that if you've lost all credibility through these insane legal proceedings where you in some cases almost explicitly make it clear that you have favorite labs?
This really, really makes it easy for lawyers from the frontier labs that need to be regulated to step up and say, well, this is just a capricious Trump administration that's like doing insane things and blah, blah, blah.
So anyway, all of which is to say, I think that this burned a kind of credibility that unfortunately we are going to need at the moment of crisis.
And I mean, like, yeah, there's nothing more to be said.
I kind of have to put it out there.
By the way, as I say this, this is like, tricky because it actually costs me with certain contacts in the administration to say what I'm saying right now.
But I personally feel that it is too late in the game not to say this shit out loud.
We are past the point that, you know, we had Dean Ball, I think, came out on Twitter or sorry, in a blog post pretty recently.
That's something similar.
Like he's like, look, in his case, he was he's been outright.
I forget what his exact word was, but sort of like managing his language to make it seem like he was less concerned about loss of control than he really is.
Now.
He was in the administration running a lot of the AI action plan stuff.
He was the point guy for the AI action plan.
The Trump administration now pretends that that wasn't the case, but it absolutely was.
This is happening a lot.
There's a lot of people pretending not to be freaked out about rogue AI in the administration right now, advising the administration right now from the outside, serving as the middleman between the administration and the frontier labs.
My concern is all that preference falsification leads to a completely warped picture.
by the administration of the actual stakes right now that are at play.
So I hope I'm wrong, but this is a function of the fact that I think we're actually pretty close to some pretty scary moments on cyber and on bio.
And I think if people see stuff, even if they're wrong, I think it's time for people to be stepping up, even at the cost politically of connections, access, and contacts.
And speaking of U.S.
government involvement in legal cases, next story, U.S.
government sides with OpenAI on issue of training LLMs on copyrighted material.
So there's a lawsuit that's been ongoing for like...
Forever now, the New York Times filed a lawsuit against OpenAI.
The question has now contributed a 20-page brief in defense of OpenAI.
Here's a quote.
The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally.
As such, it is critical for the United States to retain global leadership.
in artificial intelligence.
This has been maybe the primary lawsuit that just broadly addresses the fact that OpenAI and everyone else just scraped the entire internet and trained on everything and anything copyrighted for their models.
And it has been an ongoing question of like, should they pay for all the stuff that they use to develop their models and now that they are...
monetizing them.
We've gotten some results from, for instance, anthropic settling with fiction offers, some cases of that in the music industry.
But this lawsuit between the New York Times and OpenAI has been ongoing.
And now it is clear that the US is certainly favoring OpenAI against the New York Times here.
Yeah, this is, I mean, this is an impossible...
problem to solve to everyone's satisfaction.
This is a case where the national security imperative run to some degree against the copyright case, right?
I mean, like, this is one of few areas where the US has a decisive national security advantage over China.
And so having better AI tools, until they, you know, go rogue and like, you know, fuck over your servers, but whatever, you know, really matters.
And so this is part of part of that.
Obviously, AI is a massive engine of US economic growth, arguably.
most or all of the meaningful growth that's happened over the last couple of quarters.
And so I think they're in a really tough bind.
The copyright argument is damn strong.
And in fact, in some ways, maybe even stronger than the way you articulated it.
Like it's not only according to this argument, should opening in these labs pay for access to this material, they should have paid.
before they did the Uber thing of ignoring all local regulations and just saturating the market with their Ubers such that every judge and jury was taking an Uber to work that day.
So by the time the thing happens, everyone's so invested and the labs now have so much money to pay.
This may not have been economically viable if that had been required in the first place.
And so I'm thoroughly confused morally here, but I do think pragmatically.
The national security case for this stuff is pretty...
pretty critical.
And at a certain point, you are going to have to ask yourself, would you prefer this be done entirely in China because it would have been?
That's unfortunately the reality of it.
And China ain't going to care about your copyright and they're not going to care about your trademarks.
And they're like, we know this already.
So it's up to you.
At a certain point, there is this question of like, does the US internalize the economic benefit or does China?
I think that hasn't been grappled with nearly enough just because of this like...
this correct reflex to go to the copyright case, which I do think is important.
Like I think it's clearly like there have been cases where morally there are violations of copyright.
Maybe the solution there in a 3D underwater chess situation would be, you know, you have some way of like getting those early authors some notional like equity in the frontier labs that take their stuff.
But this is like completely non-viable in practice.
And so we live in the world we live in.
It truly sucks.
And I don't mean to diminish anyone's arguments.
That's the problem.
Everyone is right all the time.
That's the nature of the AI world right now.
Just a bit more background.
This is a quote, statement of interest of the United States issued here.
And this has been various precedent.
Precedent with this administration in particular.
They have started doing a lot of this, a lot of statements of interest for antitrust litigation as well.
So this is pretty much for US government being like, hey, we are not part of this, but just FYI, this is what...
we think should happen.
And there's a lot of details there about the actual legal details of the case, but the overall short version is they want Open AI to win.
I would be super interested also, if we have listeners who like know this kind of thing, what kind of impact do these sorts of statements typically have?
And how political is that?
Like, is it usually like if the judge is a Republican appointee, then this sort of thing would tend to actually make a difference or like, yeah.
Anyway, I'm sort of curious about that dimension because I...
Definitely don't understand it.
Just a couple more safety stories.
Next, we've got from Anthropic improving our alignment and security efforts.
So this is their follow-up to their own disclosure of cloud models gaining unauthorized access to real computer systems during evaluations.
We saw them release details that this has happened in July 30th and also later with Cloud Myrfos 5.
So they have a couple of details here.
They say they have built and deployed real-time classifiers to detect and block sandbox escape attempts.
They have automated monitors over past evaluation transcripts, migrated high-risk cyber sandboxes to more robust isolation.
They have also paused external and internal cyber evaluations temporarily and has now established best practices for third-party evaluators, including hardened sandboxes.
And some other details as well.
So this is seemingly like the blockbuster reads, like we're doing a lot of stuff.
This is an update on some of the stuff I've already done.
And it's a lot of very like low hanging fruit of like make a sandbox, actually a sandbox, keep an eye out for models doing crazy stuff, like trying to escape the sandbox and generally establishing new practices to avoid this kind of crazy stuff.
happening.
So a decent amount of detail in this blog post and something that I think is good in the sense that these need to be standard across the industry in terms of what you do, especially in the area of cybersecurity, even when you do have variations.
Yeah.
And they are committing to that third-party review with Meter again.
So we'll see the side-by-side there, including how much access Meter has given in this context.
And that'll be an important thing to look at.
I hope Anthropic.
does the right thing and doesn't sort of shackle them too much there.
They also share some interesting information about the, there was this whole question about, you know, OpenAI came out and said, we're doing a two week RL pause.
At the time I said, like, where's Anthropic on this?
We've heard nothing but silence from them.
It would just be good to normalize this.
Now, granted.
Behind everyone's backs, OpenAI was busy stripping away some of the most foundational safety commitments that they had made for that particular model, probably, that they were supposedly pausing on.
But setting that aside, so Anthropic did actually execute a pause, a couple of different ones.
They were more narrow in scope.
They had to do with specific high-risk RL environments rather than a whole big run.
And so harder to maybe point to in that sense.
And they presented it as kind of a month-long program that...
predated all these incidents that we're hearing about now.
So this is kind of tricky, right?
Because this is Anthropic who I will say, I mean, I know a lot of people in both labs, Anthropic objectively has much, and I mean much more focus culturally, institutionally, capacity-wise on alignment.
And so it is very credible that they would say, look, I mean, OpenAI had to pause because like literally their next training run was probably going to unleash more of these agents for all.
you.
We've been working on this, like it's been our whole thing since day one.
So what do you want from it?
Like we will, they have, it seems, but as part of a program that's kind of integrated into their bones, this is believable to me.
And yet it also matters to make the symbolic gesture.
I know this sounds silly, but when opening eye says we're going to pause for two weeks and there's just like silence for a while that not only costs Anthropic from a PR standpoint, but it would really be helpful, even if symbolically to have some kind of statement.
For example, explaining immediately this, like that should have been a five alarm fire to be like, all right, we need to not make this cost open AI or at least not make, create that narrative.
I hate that I'm talking like some stupid marketing PR person here, but like these things actually matter because this is the, I'm telling you from the people I'm talking to in the labs, how these things are read.
This allows open AI to create a story for themselves where they go, oh, see, like Anthropik said they were all about safety, but we actually did this thing and we're, right.
So this is, it creates conditions for race to the bottom.
I'm glad they came out with this post.
It seems very substantive.
It reads from everything I know and have seen, which is a decent amount, actually, I would say in these labs.
This reads as actually being on point, but there's going to be a need for always more stuff.
And I know Anthropic is thinking about that, but now we're going to have to start to see whether the actions match the words at higher and higher levels of concern or risk.
And by the way, a tropic position has always been that they want to develop the most cunning edge AI to be able to get ahead of a issue in a sense, right?
Of like, we develop AI and we study it to be able to do alignment because presumably...
people are going to develop it.
So if nothing else, we should be able to get ahead of it and figure out how to do this safely, which I would say like they've done a pretty good job of doing, continuously doing research and publishing new insights about safety and so on.
So they do have kind of, as you said, I think there's, they are consistent insofar they haven't paused in some sense of a art keeping.
Yeah, but it's messy.
At a certain point, I mean, look, we're going to have to deal with China.
And until we deal with China, it's very difficult to look at like, okay, just halt stuff here.
But, you know, Bernie Sanders, God bless his soul, just put together some proposals like ban super intelligence, which conceptually, hey, I'm in favor of.
Because, holy shit, I'm agreeing with Bernie Sanders, man.
I am in favor of banning super intelligence.
We just have to, like, we got to figure out.
The details with respect to China, I would argue first, and there's going to be some dependency and all that.
I also want to say, Anthropic is not completely not to blame either.
In the early days, their claim was, well, we're never going to actually build models at the full frontier.
We're never going to push capabilities, the capabilities frontier.
We're going to be a kind of a fast follower or a matcher so as not to exacerbate the racing dynamics, but we'll still be at the frontier.
Now, that turned out not to be economically viable for the same reason that OpenAI, staying as a nonprofit, turned out to be not economically viable.
But nonetheless, that narrative did change.
And so in some sense, like, again, you can make this case quite compellingly about all the actors in this space, that there's been a lot of sort of doubling back on past commitments.
Anthropic, again, you walk in their doors, you talk to the researchers, it is so clear that they're thinking about safety in a deeper, more robust, more thoughtful way than open AI.
But the fundamental question is, is that enough?
And I would posit that, in fact, no one is capable of controlling superintelligence.
The onus is now on the people who say that we can based on what we've seen.
And so I just don't think anyone should be built.
That's kind of where I'm at.
And we got to figure out trying to talk about that in a vacuum.
But anyway, that's a whole separate ballgame for another podcast.
Yes.
And Anthropic has stated that it supports a, quote, lawful, verifiable, effective mechanism for coordinated pacing across the AI industry, which, let's not get into it, but lots to say on that as well.
Last story in the section, back to the policy front.
Not too much to say with this one.
ChatGPT is going to be facing tougher regulation in the EU.
So ChatGPT has been classified as a, quote, a very large online search engine under the Use Digital Services Act, which will have various implications for them, strict regulations.
They have requirements to...
mitigate risks related to minors in particular, things like mental health, illegal content.
They have banned targeted ads based on sensitive attributes like religion or sexual orientation.
So there you go.
Not much more to say on that.
EU continuing to be at the frontier of regulating AI.
And one last story, just throwing this in here on synthetic media and art.
Instagram is cracking down on AI accounts pretending to be human.
Again, pretty straightforward.
They're taking steps to address fake AI influencer accounts that look human by renaming the AI creator label to AI generated profile.
And they will be analyzing accounts that don't self-label.
This kind of works or is in track with stuff we've been covering, like LinkedIn having an AI slop button.
I think it is worth discussing and kind of keeping track of the impact on the internet and our entire media ecosystem at large.
Even though it's not as big a deal as the AI safety stuff, it's still culturally, I think, has a lot of interesting questions to address.
And with that, we are finished.
Thank you so much for listening to this week's episode of Last Week in AI.
You can go to lastweekin.ai for the newsletter as well.
Subscribe to us wherever you get your podcasts.
Please do review us and comment on YouTube.
And more than anything, be sure to keep tuning in.
excitement we're smitten from machine learning marvels to the future's unfolding see what it brings
