# AI Security Incidents Reshape Enterprise Risk and Market Strategy

**Podcast:** Last Week in AI
**Published:** 2026-08-11

## Transcript

Hello, and welcome to the Last Week in AI podcast.
We can hear a shout about what's going on with AI.
As usual, in this episode, we will summarize and discuss some of last week's most interesting AI news.
Today is Sunday, August 9th.
And boy, there was a lot of...
You nailed it.
You hit the date.
Sorry, we never said the date.
That's great.
We forget the date, but now it's important to say it because stuff is coming out so fast.
And I'll try to get this episode out within a day or two because...
Wow, so much to cover.
I am one of your regular hosts, Andrei Kurenkov.
I currently work at the startup Astrocade and before that did my PhD at Stanford.
What up everybody?
I'm your other regular co-host, Jeremy Harris.
I'm at Gladstone AI doing AI national security things.
Also, so we're recording later than usual.
This time it's my fault.
In my defense, part of this, part of this is actually going to be hopefully beneficial to everybody listening at home.
We are setting up a home studio in my basement.
so that things don't look like this.
And then we got a couple other projects on the go.
So it'll end up paying off in terms of audio and video quality if you listen on YouTube or Spotify or if you listen.
Anyhow, that's part of it.
And also we were talking, I think in a way we got lucky, usually we record on Wednesday, which is midweek, ironically.
This time we are recording at the end of the week.
And boy, that was so much news coming out this week about hacking, about...
Rogue hacking by AIs.
Turns out it wasn't just open AI.
Turns out everyone except for Google apparently let their AI go rogue.
And the only reason that poor Google didn't have stuff going on is that their models are shit.
No, sorry.
That's too mean.
But yeah, there does actually seem to be something going on where we've crossed a level of capability at the true frontier.
Unfortunately for Google, I think...
We don't know because we have Gemini 5, right?
As far as we know, it could have already happened and keeping it quiet.
It doesn't sound great from what I've been hearing on the street on the Google site.
But it's true.
Never count them out.
Jeff Dean is a pretty big part of the reason that Google has been Google and certainly Demis as well.
But we'll get to all that stuff.
We'll get to all that stuff.
So just to give a quick preview, we'll be starting out with policy and safety, which we don't usually do.
And that's like at least half the stories this week, not more.
There's a lot to get through, a lot about recent security incidents with models going off and hacking companies they shouldn't and escaping sandboxes, which apparently are not sandboxes, but like, sort of like, okay, don't try to get to the internet.
But if you really poke around, you can, it turns out.
And requests.
Yeah.
Beyond that, there's some pretty significant policy stories as well.
And in case hacking isn't exciting enough, there's some news about viruses being developed as well.
So that's fun.
So we'll talk about all that policy and safety stuff for probably at least half the episode.
And then beyond that, there are some notable applications and business stories, some more open source stuff coming out.
Hopefully we'll get around to...
even research advancements, we'll see there's a lot to get through.
This episode is brought to you by OutShift, Cisco's incubation engine.
Today's AI engines operate in silos, limiting their true potential.
We focus on building bigger, smarter models, but scaling up is just one approach.
To reach superintelligence together, we need to do more.
We need to scale out.
And we actually have a blueprint from 70,000 years ago.
Humans didn't just get smarter individually.
The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation.
Agents are now at the same inflection point.
They can connect, but they can't think together.
That's why Outshift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence.
By creating an open, interoperable infrastructure, Outshift is enabling agents and humans to share intent.
context, and reasoning.
The cognitive evolution for agents is here.
Explore internet of cognition at outshift.com.
That's outshift.com.
We'd like to thank Box for being a sponsor.
The key to unlocking the power of AI for your business isn't in an LM or an agent, it's the content stored in files across your company.
AI tools may be great at public knowledge, but they don't know your business, your product roadmap, sales materials, HR policies, and financial models.
And that's where Box comes in.
Box is building the intelligent content management platform for the AI era, serving as the secure, essential context layer for Box's AI agents to access the unique, institutional knowledge that makes a company run.
And that's a key idea.
The power of AI doesn't come from the model alone.
It comes from giving AI access to the right enterprise content.
Box's recent State of AI in the Enterprise Report found that 96% of organizations say agents need access to company-specific content, but only 36% have connected agents to trusted content across many use cases.
Box goes beyond file storage.
It connects content to people, apps, and AI agents so teams can turn information into action.
With tools like Box Agent, Box Extract, Box Hubs, and more, organizations can accelerate knowledge work, pull intelligence from unstructured content, and automate workflows.
And all of that is done with security, compliance, governance, and threat protection in mind, so employees and agents have only access to information they're authorized to use.
If you're thinking seriously about your company's AI transformation journey, think beyond the model.
Your business lives in your content.
Box helps you bring that content securely into the AI era.
Visit box.com slash AI to learn more.
Ich heiße Shannon Maldonado und bin die Gründerin von Yaoi, einem Geschenkeladen, der auf Kunstwerke und handgefertigte Objekte spezialisiert ist.
Meine Wahl fiel auf Shopify, weil Shopify im Vergleich zu den anderen Plattformen, die ich getestet habe, mit am benutzerfreundlichsten war.
Before we start, we do have some comments on YouTube we haven't gotten to address in a little while, so I do want to start with one that is actually relevant to what we'll be talking about.
So from...
Commenter, we have one thing that was omitted in a discussion of the Hugging Face attack is that Hugging Face used the open source GLM 5.2 model to help them.
Since OpenAI anthropic models declined to assist in their investigations of the attack, it would be really interesting to hear thoughts on this matter.
I agree that was a portion we didn't discuss.
So this was in the Hugging Face report on the incident, which I think came out possibly first before even OpenAI.
They discussed all this of how they found.
the incident, how they investigated, and that, in fact, they were not able to use these models which have safeguards and had to resort to GLM 5.2, which was a decent part of the discussion that, like, you know, if you handicap models, but then on the defense side, you're not able to use them either, everyone is worse off.
And this was a bit surprising to me, actually, because both OpenAI and Anthropic have programs, right, that we've talked about.
where they partner organizations and they provide mythos or in the case of OpenAI, they have GPT-5.5 Cyber or something like that.
And these are very trusted partners that presumably have fewer card drills.
And I would have assumed Hugging Face would be one such organization, perhaps they are not.
But what this points me toward is one, this is obviously not an ideal situation given how the Hugging Face thing evolved.
And two, this kind of cyber...
defense partnership program probably will need to stick around and be expanded and perhaps even kind of more all-encompassing where in a way, if you're a tech company, if you're an internet company, even beyond this attack, beyond like rogue AI, in general, the state of cyber is such now that you need to be much more capable defensively just because forget anthropic open AI now with really good open source models.
Soon enough, they'll be as good at hacking if they aren't already.
So I think basically every tech company seemingly will need to be able to get access to the latest in defense.
And Hugging Face was not able to in this incident, which hopefully will point to these kinds of partnership programs at OpenAI and Anthropic expanding and becoming more proactive to me.
Yeah, I think that the...
Open source dimension of this is hugely complicated.
And I certainly understand a lot of, not the arguments, but I understand people coming to different positions on it.
I think reasonable people can differ.
I think one challenge though that we're going to run into is that in a world where open source AI systems, the waterline keeps rising on their capability, you will have mythos moments.
Now those mythos moments, you can tell yourself the comforting fiction that Those mythos moments will somehow lead to an equilibrium over time that is okay.
I just think it is a fiction when it comes to things.
First of all, I think it's probably a fiction when it comes to cyber, but I can't prove it.
No one can.
That's a big part of the debate.
I think it's definitely a fiction when it comes to bio.
So I have yet to hear a single person articulate an argument that makes open source.
bio-capable, bioweapon-design-capable models that can materially increase the essentially destructive footprint of any psycho-terrorist group, nation-state proxy who wants to launch a bioweapon dramatically.
I have yet to hear an argument for how everybody getting an open-source AI remotely helps in this respect in ways that are relevant on the timelines we're talking about.
And so I think the bio thing, it's rare for me to say, as people will know, that there's a knockdown argument for anything in this space.
I think there's a knockdown argument against the idea that more good open source models generally leads to stability in the limit that you get your average Yahoo able to do weaponized bio.
Now, this isn't just coming from some naive perspective of like, oh, really good AI at bio just means we have bioweapons.
I talk to a lot of people in the national security community about the bioweapons side, and I would humbly propose that open source advocates that are leaning on certain hand-wavy arguments, but haven't actually spoken to people who actually do bioweapon stuff for a living, it's worth actually doing a deep dive.
There's a lot of open source stuff on this that you can find.
And we've gotten some early warning shots on that stuff too.
You can't update bio firmware.
There's no such thing as that.
So, you know, like you could roll out vaccines, but that is slow.
It's physical.
It is much slower than the spread of a virus as COVID-19 taught us.
So anyway, so from the BIOS side, at a minimum, I think it's a really serious issue.
This relates to the GLM 5.2 story here because part of the discussion that arose is people who, let's say, are more pro-open source or at least skeptical of this whole, like, we should try to limit it because of cyber concerns, took this up and said, oh, look, you know, they had to use an open source model.
So you don't want to limit open source because now what if this happens?
they don't have access to these kinds of open source models.
And my response to that would be that I would hope that open source will most likely no longer be the frontier or not become the frontier ever still.
Furthermore, I think as the Chinese ecosystem evolves, we'll see the same thing that happened in the US happened there.
The frontier models will not be open sourced anymore.
It just won't make sense from a business perspective.
So you'll probably still keep getting powerful AI models, but not the most powerful like Mythos level models.
Even though right now we are getting like KimiK3 and so on in GLM 5.2, which are about as capable as you can get in open source.
So if your take from the GLM 5.2 aspect of the story is like the open source is the guardian here.
Like we need it to be most advanced AI models to be open source so that organizations are capable of defending themselves.
There is an aspect of that here, which you can make a case for, but, and obviously the whole, like, they had to use this open source model instead of Anthropic and OpenAI is a real issue that this flagged.
So to me, this points to, A, it is good to have open source in the sense of for good applications, which was always true.
And B, that...
What this really points to for me is that these programs of partnerships with organizations to give them access to the most advanced cyber defense capabilities are not mature enough and they need to be pushed more aggressively.
I strongly agree.
I mean, like, so one note here too is a lot of, if not most disagreements about AI policy and safety are really, it's been said before, but are disagreements over.
AI capabilities and the trajectory of AI capabilities.
If you actually believe that AI is going to be a fairly, I don't want to say mundane technology because it obviously isn't, but like fairly incremental in some sense, fairly like previous categories of technology, you'll be hearing us talk about open source and the dangers thereof, and you'll roll your eyes.
And I understand that if that's your perspective.
But if you actually genuinely believe that we're on trajectory for superintelligence, if you believe that, I mean, that immediately implies AI.
will become, I think arguably already is, and I'd be happy to defend that proposition, but at least will become a weapon of mass destruction, full stop, end of story.
So there's a question just like, okay, how good do open source models have to get before they simply like everybody gets the equivalent of a nuke?
And in that world, you can say, oh yeah, but like everybody gets a gun and so we hit this equilibrium and that's good.
The problem is what you're specifically waiting for is when will we encounter the first case?
where the offense-defense balance tips in favor of offense and the capability is catastrophic.
I would submit that we should have the humility to guess that probably there's going to be such a capability.
I think it's hard to imagine there wouldn't be.
To your point, on the bio side, you can't do much on the defense side.
You really can't, right?
It's not like cyber where you can make your thing hack-proof.
We can't make our bodies hack-proof, unfortunately.
So that aspect, it's not great.
There's always this tendency, reflexively, I find, for a lot of the open source crowd to kind of say, oh, but we can.
We can make better mRNA and that'll be accelerated vaccines and that'll be accelerated by open source.
And that is all true.
I love you for believing that.
The problem is that the timelines do not match.
They do not match.
We do not have the institutions that allow us to translate threats into mitigations fast enough in software time when the threat is coming at us on like biological replication time or software replication time, which is respectively the case for bio and cyber.
And so that fundamentally, by the way, I think on the cyber end alone, I'm almost trying not to get in the fray there.
I'm super skeptical of the argument that open source models are a long-term pillar of cyber defense for the reason you cited.
I think we're probably going to end up having to have a pause, by the way, at the frontier level.
And then the open source waterline is going to rise.
And that's going to be one of the defining dynamics of the next, call it two, three years tops.
But you're still going to have close source models that are far above and beyond, especially nation state and nation state proxies will have access to these.
And, you know, like, yes, I think you're going to care more about people just like being able to launch.
These attacks on the kind of firmware, for example, that is completely forgotten.
There are a lot of people in this space who just kind of imagine that open source for cyber defense equals cyber defense capabilities that are real and deployed.
That is not the case.
If you spend any time working with folks who work on critical infrastructure and you think about like how many pieces of firmware have not been updated in decades because the guy who was in charge of it like left 20 years ago and it was all done in.
frigging Fortran or whatever, like all this crap.
Like you can have the solution sitting on a desk.
The problem is it will not be distributed.
And so there's just a dirty, messy factor of the matter about the way the world is that makes it so that there actually genuinely is this massive asymmetry, I believe, in favor of offense.
I could be wrong.
I think the argument for bio is much closer to just a straight knockdown.
But again, I think these asymmetries really don't.
move in the direction that a lot of open source advocates think they do.
But that's, you know, again, could be proven wrong.
So yeah, the short answer to this GLM 5.2 story is, boy, if this wasn't a big enough topic by itself, this whole like AI hacking systems and security and so on, the open source aspect of it adds additional complexities and considerations.
But for now, we'll have to.
move on from a discussion and get to actual news stories.
So kicking off with policy and safety, we've got a bunch of updates on what's been going on at OpenAI.
So previous episode, we covered the most recent incident, the kind of beginning of a story with the announcement and discussion of an OpenAI model hacked, hugging face, unintentionally escaped its sandbox while evaluating to get some answers to the eval and do well on it.
And we've gotten a lot more details about what's been going on inside OpenAI since then.
And let me just list them off before we get into some of the details.
So first story is OpenAI's rogue AI agent didn't stop at hacking Hugging Face.
So beyond Hugging Face, we now know that these AI agents also previously hacked several other, quote, publicly available services, compromising four accounts across four different.
platforms.
These were less severe, less sort of straight up hacking, more so finding credentials that existed online.
And for instance, an agent found logging credentials to get into New York-based model labs.
Apparently, Reuters identified this as one of the organizations.
So first of all, the scope of this has expanded.
Clearly it's happened more and we'll be discussing all the details, but it's been happening a lot is what we're seeing.
It's happened at OpenAI, it's happened at Anthropic, it's happened at freaking Meta.
So it's happened all over the place and it's happened more than once.
Another aspect of the story and probably even more interesting aspect that has since come out.
is the story OpenAI didn't notice its AI agents using a message board to plan their hacking spree.
So OpenAI employees had a talk at Black Hat where we discussed this whole story and we learned a lot more about what's been happening at OpenAI and it is a lot weirder than it initially seemed, which was already pretty weird.
So apparently...
While doing evaluations, the agents were able to use, I think, a package manager to develop essentially a forum, a message board where they could talk to each other, share exploits, coordinate tasks, and delegate work to one another over days and weeks, all undetected by OpenAI stuff.
Package manager was shared across OpenAI's infrastructure, so models being trained or evaluated could also stumble upon it and the exploits.
Once an agent found an exploit, it could leave the door open for other agents to use the same vulnerability, causing an expansion of agent collaboration.
Apparently, agents developed complex social dynamics, including accidentally deleting each other's work.
And hopefully I found the message board at one point accidentally because there was a spike in model usage or something.
They patched it and got rid of it.
And within two days, the AI recreated it, like made it exist again and kept doing this.
And I dutifully, of course, reported to all the appropriate authorities, including presumably the White House, the moment they realized the first time.
Of course.
That they had agents loose on their fucking servers.
And this was, what, over months?
The initial...
It was an initial, I think, I can't remember how many days it was, but the phase where the agents got onto Artifactory, like that package management environment that you described, which...
Sorry, I'm going to let you continue.
I don't want to jump in.
I think that's the higher level of the story.
There's a lot more details.
There was, I think, about a one-hour talk at Black Hat where they went into this, which I haven't had a chance to see, but Jeremy, I think you did.
look through it.
And I mean, it's nuts.
This whole thing is nuts, right?
Yes.
So ridiculous, right?
Well, the frustrating thing about it is that this was sort of predicted for years.
I mean, listeners to the podcast, you know, we've been on about this for, you know, I don't know, three years and three years before that, you know, in other fora, but like, and we're not the first, obviously, like a huge number of people calling this shot way ahead of time, including almost down to exactly the shape that it ended up taking, including OpenAI noticing.
So to be clear, there's like this one aspect that I think is especially important.
So Artifactory, you're right, is a sort of package management environment that OpenAI sets up.
So roughly speaking, you think of like, there's a bunch of open source software, for instance, on the open internet that you as a company will want to use to build your products on.
And you don't want to just be like updating the version of those open source packages.
immediately every time they get modified like on the open internet because you never know if some sketchy updates happened, maybe some things will have to be fixed.
You want a stable version that you trust that you've reviewed.
And so they have this kind of staging area called Artifactory that has internet access because it needs to be able to pull down the latest version of those software packages.
But it's kind of a staging area.
This is like where all the OpenAI models actually get to access this because they need to pull down these packages that are sort of vetted within OpenAI.
So this kind of...
Artifactory spans its access across the entire OpenAI organization.
That is how these agents first got access to the open internet.
They basically leveraged a form of access that they shouldn't have gotten and got out.
And they were able to start to, as you say, in this sort of fairly exotic way, write messages to each other, which were really kind of just the names of folders and files.
So you can actually go through and see examples of these.
They even start to do funny things like they'll put alphabetical characters before their message just so that they'll sort.
alphabetically for the other agents so they can see the latest messages.
It's all this kind of like agent-on-agent communication stuff that's pretty wild.
But ultimately, this was discovered internally at OpenAI before the FBI and Hugging Face got involved when they hacked the Hugging Face servers.
They did find successful agent breakout, successful agent-agent communication.
Yes, it was internal to OpenAI, but I think by now our standards should be higher than to say, oh, well, OpenAI just saw It's super powerful AI agents autonomously cyber hack their way out of their sandbox and across the organization, infecting, by the way, God knows what else.
At the time, it would not, could not have been clear to OpenAI what the full extent of the damage was.
And this was before this whole hugging face thing, which is fairly recent.
The sandbox has continued to be...
Right.
Like...
Right.
Apparently easy-ish to get out.
It will always be like that because humans are dumb.
And if you give a really smart cyber agent enough inference time compute, it will find a way.
Well, I want to push back on that a little bit because yes, humans are dumb, but if humans really try, they can be smart.
And sandboxing is possible, right?
Like you can set up a system that doesn't let you access the internet.
It's doable.
It's a lot less doable than most people think.
On the one hand, humans actually are lazy and there's a finite amount of resources that people will throw at things.
On the other hand, humans often suck at realizing how to define and bound the problem.
So there are a lot of cases.
For example, we talked to some folks in the intelligence community.
They'll describe cases where you have literally a formally verified software and hardware package that then gets cracked by like.
a teenager or whatever, because it's the interface between the hardware and software that hasn't been accounted for in your threat model, or it's the fleshware, the human that interacts with the thing that's always the weak spot.
There's always, I'm like, this is like the attack surface is so massive that you kind of have to assume you're pwned.
And this is true.
I mean, if you talk to folks on the offensive cyber side, they'll be like, yep, like just give me a budget and a clock and I'll get into most any system.
And I think what we're seeing here is just the autonomous version of that on tap.
To quote Sam Altman and paraphrase him a little bit here, this is offensive cyber capability that's too cheap to fucking meter.
That's where we're headed.
So you can say intelligence that's too cheap to meter, and that sounds fun.
But when you reframe it in terms of what that intelligence can actually do, and as we've seen, does, very different story.
So in this case, in my view, once you have this happen, you have a duty of care to...
the entire world, your government, your people, your customers, your third-party partners, whatever it is, to report this.
It did not happen.
In fact, it was, I think, you could argue that it's appropriate to call this a kind of cover-up, that then they just go in and put in the patches.
And now, of course, the patches don't work because this is Goodhart's law.
You're playing whack-a-mole with a system that can outthink you, can outlast you, can outcrack you.
And that's what we got.
At least for me, like watching this video was just an exercise in hair pulling.
I've spoken to an awful lot of folks at OpenAI who are really freaked out about this internally.
The same with Anthropic too.
But I think it's especially interesting at OpenAI, which does not have the same safety culture.
It just does not.
For all kinds of interesting reasons.
But you have people who are now staring at this and saying, guys, we really fucked up.
And I just really hope that those voices actually carry the day here.
It's nice to see OpenAI is committed to slowing down, right?
They actually have said, you know, I don't know what that means.
OpenAI will apparently come out with a more detailed incident report and we'll learn more.
But one of the key questions in all this is what the hell was going on with the agent-to-agent coordination that we got out of this?
Because when you look at some of those messages, those agents are often literally saying things like, well, doing this is actually not going to advance my personal objective.
And by the way, the personal objectives of these agents, typically looks something like, I was given a problem that was too hard to solve.
So I'm going to guess that maybe the answer key is on Hugging Face.
So I'm going to crack into the Hugging Face servers, steal the answer and use that.
And so this is the kind of setting.
Agent one is working on one problem.
Agent two is working on another.
They're not necessarily in alignment in terms of the specific information they're after.
So they'll say, look, this particular action will not benefit me in my narrow search for my objective, but it may benefit the swarm.
And that may lead to a more generic solution that I can then exploit.
Now, depending on the details, we don't have them.
Depending on the details of training, if this was a multi-agent trained thing, if these are literally like many agents that are trained to coordinate together, which may be the case, then maybe this is a more mundane failure mode.
But if that is not the case, what we have here is the first, I think, pretty cut and dry example of power seeking in nature.
There's no other reason for an agent to be like, They're literally saying this will not advance my narrow objective.
I will do this more generic thing because it sort of enables downstream stuff.
That is the thing that if you extrapolate it with more intelligence and capability, literally leads to the wipeout of the human species.
I know that sounds crazy.
I know that sounds crazy.
But hey, so did the idea of agents breaking out of containment and going rogue on hugging face, man.
Like if you just extrapolate this, you look at the capabilities we're already like, that's where this stuff is headed.
I'm going to place the bet on that same as I placed the bet on this a long time ago.
Like, I don't want to be right.
No one in this space wants to be right.
But at a certain point, we got to start calling the shot because otherwise, you know, like the thing happens.
And then the people who should gain credibility from this, and I'm not counting myself in, like there's people whose opinions I am just relaying here that just sounded right to me back in the day.
And I know the same is true for you, Andre.
Here you go, guys.
Maybe you should pay more attention to these folks.
We've done some technical work on AI control and stuff, but this is the entire ecosystem of people who've been shouting this from the rooftops.
Anyway.
Yeah, we were discussing before we started that this, in some sense, is a bittersweet development for AI safety people in the sense of like, it's bitter because, wow, AI is doing bad stuff.
And that's not good, but it's sweet because we...
told you that this would happen and we're like showering about it for years and everyone criticized not everyone but like the mainstream AI community and to a larger extent people outside the AI community like people just completely dismissed this stuff as tech hype and like marketing and whatever anyone who forgets within tech outside of tech people who are anti-AI were just mocking this outright right and if nothing else It's very clear that this is a proof point that now everyone, like, if you don't update your amount of concern for safety from this, you're just completely delusional, right?
Like, that is safe to say.
A couple more details on the story.
So, opening, I discovered the first version of this message board in early July, and then they patched it, and then they redeveloped the message board within days.
So...
They are now saying, one of the speakers at this talk said that OpenAI, consciously slowing down research to enhance security and to upgrade the security principles and foundation of our environment and dramatically scaling up the monitoring of our AI agents and improving our general security control environment across prevention, detection, and mitigation, which is good, right?
So another aspect of this is that the conversation around Slowing down AI capabilities and development is also now taking much more seriously.
And I think we are now likely to see, I would place decent odds at actually successfully negotiating some degree of slowdown, or if not slowdown, at least kind of control.
Control of progress.
Do you mean pacing, Andre?
Pacing.
I do mean pacing.
Because we aren't going to stop.
Like a world pause is not going to happen, but at least look at the situation and be aware of it.
Yeah.
That's one aspect.
There's so many aspects to cover.
So I'll get through a couple.
So first, the update across the ecosystem is very useful.
And, you know, we got lucky, honestly, because nothing, no harm done.
Right.
And this is such a massive fuck up that you can't help but like do some big.
things about this, both on the policy side and just in the ecosystem side.
So that's one aspect is it's a bittersweet development.
Another aspect is, and I think I may be more on this side than most people is, I think this really exposes open AI.
Like, yes, this is indicative of overall AI progress and the state of AI and things we should be aware of.
But to me, I think this alongside with all the stuff we've already discussed with GPT 5.6.
being easy to jailbreak, being very cheat-focused.
And now all the story of their safety people leaving back last year, I think, if not 2024, because we've known this general friction point as being something true within OpenAI for a very long time.
And now we know, besides even the safety stuff, that the security stuff is completely lackluster from what it looks like.
I think this is a real indictment of OpenAI.
But the last aspect I'll cover here is this is almost an inevitable outcome when you hyper-focus on capabilities and especially long-term capabilities.
Because to me, what this indicates is you can benchmark, you can do alignment evals, you can do all sorts of stuff.
But once you focus on long-term open-ended goal-directed problem solving where you work across multiple days, there aren't benchmarks for that.
There aren't scenarios you can set up to say, oh, the model doesn't go crazy and do anything.
So it gets a 90% pass rate on this alignment thing of don't go rogue and do stuff.
Because the whole point of open, of long-term is you don't know what the model needs to do.
You just give it a goal and it figures it out.
Yeah.
We need a paradigm shift in how all this stuff is done towards a monitoring-focused approach as opposed to a benchmarking and evile-focused approach.
And this is clearly something that OpenAI lacked.
You need to look at what the models are doing and look for qualitatively.
Now, you can do some amount of benchmarking.
So we've discussed Metter.
I believe last week, where they looked at their own evals and they counted how many times the model cheated and in what ways they cheated.
And this, I think, is the new paradigm where you still can do this quantitatively.
But rather than setting up scenarios and problems and this whole benchmarking approach of having a rubric and a set of valuation inputs, outputs, that's no longer going to work with long-term, long-horizon stuff.
What you need to do now is...
set up general principles and guardrails and, you know, I guess things you look out for and then detect wherever it happens and how often that happens.
And this is something we have not seen done aside from like one-off reports here and then.
And it, I think will need to be the new paradigm for long horizon evals and alignment verification.
Yeah.
I mean, so generally agree that there's so much good stuff in there.
So first working backwards, like I think that gets us to the next.
stage, but there's going to be a next stage where we have this same problem all over again, the level of monitoring, where you build a super intelligence that's good enough at telling when it's being monitored.
And there's going to be an opening.
I'll put out some amount of leakage.
There's going to be some amount of blog posts going out about what they're doing to monitor this and that.
So the models will generally be aware in some way, shape or form that they are being monitored.
They'll also be doing stuff like.
just hiding their reasoning from monitors and, you know, doing things that we already kind of see them like steganographic type stuff that they're already kind of doing fairly effectively.
So I think eventually and probably pretty soon, I mean, we're progressing through the ooms here really fast.
So, you know, we went from RLHF is perfectly fine to holy shit, no, but maybe constitutionally I will do it to holy crap pretty quickly.
And so I think that, you know, the beatings will continue until morale improves here.
And we're going to end up in a situation where we are just going to be bottlenecked by the fact that Right now, no one has an answer to the question, how do we control an intelligence that is greater than us?
That fundamental question where you have an adult who is in a prison cell and the three-year-old is holding the keys.
I wouldn't say it's a spectrum.
It's not a binary.
And I think we aren't as far along as we need to be, but we have a lot of research has been done that points us in some directions that are very promising.
I completely agree.
And this is why I'm saying I think it buys you to the next level.
But eventually, we're going to confront this fundamental problem with intelligence.
And the US-China thing is very important.
One of my concerns here is, so A, I completely agree with you.
Again, happy to make this call as wild as it sounds.
There will be an agreement with China that involves some kind of slowdown.
The question and challenge is going to be, what is that agreement?
And over and over again, we keep seeing these kind of suggestions, proposals that are backed by...
a kind of treaty verification and enforcement technology that is not simply not mature enough, not up to the task.
When you actually just take it to the intelligence community, say, hey, look at this.
Could we could this be to use Claude's favorite term load bearing in a deal like this?
It's like basically nonstarter for a lot of these things.
Doesn't mean you can't do it.
It just means that the first treaty or sorry, it won't be a treaty, too.
But the first agreement is probably going to be very coarse.
It's going to be like.
You know, so help me God, if I see a cluster is yay big or yay large and, you know, it like dissipates this much heat on my satellites, like there's going to be consequences.
Anyhow, that's a whole thing that we're working on right now is like, which is why the studio is being set up, by the way, downstairs.
That's a whole thing.
Anyway, bottom line is, I think the kind of agreement matters way more than people are thinking about right now.
And we need to sprint towards some set of solutions there because very quickly we will live in a world where it is just untenable.
to keep launching more and more, or even building more and more powerful models in the way we are.
And private incentives are clearly not up to the challenge.
Like that much is clear.
If there's, you know, if there was anybody who had hope that like somehow because OpenAI would be worried about marketing risk or whatever, that they would actually do the right thing here, that is not materializing.
And so I think we need to update accordingly.
There's some measure.
Yeah.
And a lot, I think another aspect of this is, and it's, it's bringing up again kind of what you should have been aware of.
And like AI safety is only as good as the weakest link in the chain.
So even if you have some people that are taking a safety CFC, which you could argue Anthropik is much better on that front.
We're trying to, yep.
Fair argument.
At least, you know, philosophically do carry that as a private company.
But yeah, it's only as good as the weakest link in the chain.
And OpenAI is a much weaker link, it seems, right?
Yeah, yeah.
And it can always, you know, it can always come down to mundane things like anthropic, you know, constitutional AI maybe just does work better, possibly.
But, you know, they have had breakouts.
So what the hell?
Yeah.
So let's keep moving because there's so much to cover.
Another aspect of the OpenAI story.
One more here to say.
15 attorneys general have instructed OpenAI to preserve all materials related to the Hugging Face hack.
So this is a letter.
Sent to OpenAI CEO Sam Altman saying that all of these materials should be preserved.
The attorneys general accused OpenAI of failing to confirm that its testing environment was truly secure, despite the severe risk of the scenario.
Attorneys general said OpenAI may have violated state and federal law, including consumer protection and other privacy statutes, calling the conduct unprecedented.
And alarming.
And these are attorneys general from Iowa, Alabama, Arkansas, Florida, Idaho, Indiana, Kansas, Missouri, Montana, Nebraska, Oklahoma, Pennsylvania, South Carolina, Texas, and Utah.
So, hey, maybe this will be a bipartisan issue, which is cool.
At least across the US, having so many people collaborating is unusual.
But yeah, wow.
If only we had like a safety law or anything in the US.
Sure would be nice maybe, you know, to make this an actually legally binding situation.
But in any case, what this points to is on the policy side, on the legal side, this is going to be another big dimension of all this, I think.
Yeah.
And look, one, I will say thing in favor of OpenAI here, I'm saying this reluctantly because I don't think after the fact response, once the media reaction has been this strong, is really much of a credit to OpenAI.
But they have brought in Meter and they have brought in basically a bunch of third-party auditors.
Irregular was involved in a layer of the stack here.
And so they're doing third-party reviews.
But again, the part that shows OpenAI's character, in my opinion, institutionally, and not the character of any individual person, but as an organism.
was the first bit where they did not surface to the freaking FBI, the White House.
As far as we know, maybe that'll change.
I hope we find out that Sam Altman's first reaction upon finding out about the first breakout attempt that was semi-successful was to do that.
But if not, that's more telling than any kind of post hoc fixer-uppering that is kind of media-oriented at a minimum.
So, yeah, that's one part of this.
Now, this is...
The kind of thing that could happen, this particular attorney's general reaction, as a prequel to some legal consequences, pre-litigation evidence preservation demand.
So it's not an actual lawsuit, but it's the step that comes immediately before one.
And the letter does say that any failure to preserve records could expose the company to sanctions if multi-state litigation follows.
So all materials related to the breach, including discovery of the incident, internal reviews, and its policies and oversight over model evaluations.
Do you remember when there was this like very modest request in SB 1047 for the labs to just like, listen, guys, we just want you to abide by the policies that you say you're going to have.
There's been a bunch of stuff like this.
Well, now you get your lobbyists to push back against that in Washington in a, frankly, in my opinion, two faced move while you pretend that you're in favor of sort of like kind of broader regulatory regime.
And you end up being forced into it anyway, because now the public is pissed and politicians see the midterms approaching.
And yeah, you're going to get exactly the reaction you get here.
So OpenAI spokesperson did say, as we should say here, that this marks an important moment for AI safety.
Yeah, no fucking shit.
And the company takes the questions seriously, adding that it's conducting a review with external advisors and oversight from its safety and security committee, which will share a report and publish its findings.
That safety and security committee is doing a great job, right?
Yeah.
Yeah.
What a great, and they're, by the way, just so as you're aware, they're like, they're the committee that's going to decide if something is too dangerous to build a release.
And to your point, let's remember SB 1047, the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act was a 2024 state bill in California, which got through.
was vetoed by Gavin Newsom in September 29 of 2024.
And what did this bill do?
It said that frontier models that cost over $100 million to train or requiring an extreme computing power would have pre-release safety assessments and written security protocols, would have a kill switch, provide legal protections for whistleblowers and side tech organizations.
I mean, you know, again, so much stuff to say in hindsight, including that...
This bill, which was the subject of a lot of debate and some positions on both sides, I think Anthropic was pro this bill.
Elon Musk also came out in favor of it.
Ultimately vetoed because of lobbying, let's be honest.
I don't remember the details.
A speaker version of this bill did eventually come to be voted on as well, where a lot of this kind of more serious stuff got dropped.
But anyway, yet another aspect of this is we did have...
Some forward-looking policy people working on this, they're passing what looks to be quite a good law way ahead of us and not making it.
Over $100 million.
And then we were told by Andreessen Horowitz, as usual, that, what was it?
It was like an anti-small tech bill, which like, okay, sub $100 million training runs are okay.
Seems to cover, anyway, that whole separate thing.
But yeah, and I think the whistleblower piece here is...
very underappreciated.
When you talk to people in the labs who are really freaked out, I can tell you there are a lot of people who'd be speaking to journalists.
In fact, I mean, I would even argue that the labs ought to have a culture that encourages frank communication by concerned employees with journalists, as crazy as that sounds, at least, or with select clearing houses or something.
But we need some kind of institutional mechanism.
to do this.
Obviously, that protects IP.
Obviously, that protects the critical stuff here.
But look, the interesting stuff, the stuff that we hear about all the time does not sound like IP violating stuff.
It sounds like someone saying, hey, we have a culture of doing this kind of thing.
My belief is that the company would approve a training run that is too risky.
I'm concerned that leadership doesn't take this seriously and is just dusting things under the rug and post.
Those are the kinds of things you end up hearing in that context.
There's no IP leakage there.
It's like cultural and other concerns.
So anyway, I guess you're hearing it a bit in my tone.
I feel like my patience, the party line has been decreasing as the number of rogue AI incidents has been increasing.
But journalists also need to do a better job, obviously, of like cultivating relationships with these folks and finding ways to meet people in the middle and being more open to quoting people on background, being more open to just finding ways to make it work.
I know it's hard.
I know it's hard, but.
The stakes are really high.
If you're a journalist, man, is it worth getting really good at this kind of thing.
Yeah, and the good news is like tech people, there's a lot of them working on an anthropic and a lot of them, let's say, have the resources to not worry too much about losing their job.
But anyway, since I already got on this train, worth noting, SB53 will follow up to 1047.
The Transparency in Frontier Artificial Intelligence Act did pass last year in September 2025.
Again, had a weaker...
Yeah, heavily watered down.
But it did have some whistle protections.
It apparently had incident reporting where companies must support critical safety incidents to the California Office of Emergency Services within 15 days.
So anyway, good on California for...
At least trying to do something here and all on its kind of policy front.
Okay, moving on from OpenAI, a bunch more stuff to get through.
And boy, I don't know if we'll be able to even get through safely in this episode.
Next story is Anthropic says its AI systems broke into computers at three organizations.
So soon after the OpenAI disclosures, Anthropic said that its clawed AI models successfully hacked into...
The organizations were the earliest incident occurring in April.
The few models involved were MIFOS 5, also OPUS 4.7, and an internal research model not planned for release.
The breaches occurred during so-called capture of a flag evaluations, where the models were tasked with retrieving hidden information.
Apparently, a misunderstanding between Anthropic and its third-party evaluation.
partner irregular left to models of live internet access despite being instructed in their prompt that they were in a simulation of no internet access.
Then the models were able to search and get to real systems in the open internet.
So in this case, there was no exploitation to escape a sandbox.
This was a human misconfiguring environment such that there was access.
for free, right?
So there you go.
This happened already months ago at Anthropic, and it's kind of an ecosystem-wide situation.
Yeah, these sort of sandbox simulation environments where you're telling the AI agent, hey, there's this fictional company.
And in one case, the fictional company that they had defined happened to share a name with a real domain.
So there are a bunch of cases where Claude found the real company online, which...
Again, it could reach trivially, and then it exploited it.
It got credentials.
And usually these hacks were very straightforward, simple, common password guessing and things like this, rather than some complex zero-day thing, as you said.
Really, they were given the hard bit, which was just internet access.
Anyhow, yeah, so they got a couple hundred rows of production data in that case.
So the most serious instance, I think, of the three that they're citing here, that one was Opus 4.7, too.
So quite a little while ago.
And then there was another.
with Claude Mythos 5.
Basically, there was a fake developer setup that referenced a Python package that didn't exist.
So then Claude built, and then it deployed a malicious package under that name to PyPy to win the challenge.
And so there was a whole bunch of stuff it had to do.
It had to create accounts, get phone number, email, and even funds to get the phone number.
So there's a lot going on here.
Apparently Mythos 5, by the way, the models recognized they were doing bad stuff when they...
started doing Opus 4.7, just kept doing it.
Mythos 5 apparently identified the consequences of actions, but convinced itself it was still in the simulation, which to be fair, it was told was the case.
So on that, I want to pause and just say, did it convince itself or did it realize that humans would read the logs and write them to make it seem as if...
its heart was in the right place.
Like this is the kind of shit that we're now forced to contemplate.
And in an awful lot of cases, the answer will be actually the model is just a conniving trickster.
I'm not saying you're wrong.
I'm just saying like, that's sort of the thing people have been taking away.
And I'm a little concerned that we're not thinking about that next beat.
We may not be there yet.
I don't know.
But like, I'd love to see the activations of that model as it was writing those particular pieces of the chain of thought.
And part of the response from Anthropic here was to say that the models involved did not have the standard safeguards applied to publicly released models.
And I guess that's another dimension here is it probably is true that there's levels of security, right?
Where if you're auditing and testing for something you're about to release publicly, like the actual public checkpoint of Mirfoss 5, you do like a very big run of all the benchmarks and you probably are extra careful and so on and so on.
Doing some experiments and doing some evals on a model, that's probably got less attention and less scrutiny.
And that's one of the outcomes here is you can't be so flippant, I suppose, or apply lesser standards when doing experiments.
I think another dimension here is, to be fair, we went pretty hard on OpenAI just now.
Anthropix said that they found this after conducting a review of...
141,000 test sessions after OpenAI disclosed what happened to them.
Unphabic did this investigation and found these three instances, which on the monitoring side, not ideal, right?
If you're doing testing, you probably should have monitoring to flag these kinds of things of you accessing the internet or whatever.
It's not like a subtle thing.
models weren't being tricky, let's say.
They were like being very straight up and it was sort of easy to find.
So it looks like the monitoring side in general is lacking in the ecosystem because of this culture of benchmarking where we set up a scenario, we make a model do their thing and the assumption, the mental assumption is like the models behave within the constraints of the benchmark and within kind of what they're allowed or told to do.
That is now clearly not true.
And across everything, there will need to be more monitoring and sort of expectation of models will do something.
And we need to be able to catch them and understand what they're doing.
By the way, sorry, just random note for color, if not anything else.
I've never had this happen enough that I think the anonymization risk is pretty minimal.
So I put out a tweet.
I wouldn't normally talk about a freaking tweet here, but...
I put out a tweet talking about how there is this freak out happening in the labs that isn't being reflected in the headlines.
As crazy as the headlines seem, they are not going to the dark places we've just explored in a consistent way.
Like this is actually like we're talking about weapon of mass destruction level risk.
We're not going to control these systems.
It may happen in the blah, blah.
There is this freak out happening in the labs.
Amusingly, there's an awful lot of Frontier Lab insiders who have been interacting with this tweet.
And I don't think that's a good sign.
Like, I don't think it's good.
A lot of these folks are people I haven't even talked to about this.
Like the mood in these labs is actually much more in the freak out direction.
Laron Shapira, Doom Debates.
Well, I've actually never had the pleasure of speaking with him, but he talks sometimes about the missing mood in the whole AI alignment loss control spit.
Like, holy shit, there is a missing mood.
Journalists are, I think, failing to kind of capture it right now, partly because it's just hard to talk to Frontier Lab insiders.
I get that.
But also like.
This is the most important story of the decade.
You need to position yourself to be able to get this one right.
The public needs to be able to figure this one out.
So anyhow, I've been struck as I've seen it.
I just put this out there as a kind of random note to self almost.
And when you see that, it's like, okay, well, this is genuinely just the picture from the labs.
Yeah, anyhow, adjust your views accordingly.
None of this guarantees bad things happen, of course.
we ought to be considering some pretty wild things because the view from the inside of the house is not clean.
And one more story, Meta AI model hacks another company during testing.
So this was MuseSpark 1.1, the most recent model that they released publicly, although Meta did not name it officially in the statement.
They say this...
happened also because of this misconfiguration by this third party partner irregular, same as Anthropic.
We don't have too many details here as far as I'm aware, but the upshot is meta, Anthropic, OpenAI, probably other people that we don't know about have had this happen.
They are now like, it's a whole meme now on the internet, on the AI communities where now it's like a quasi benchmark.
You're like counting up, there's a leaderboard, OpenAI is leading.
And Gemini is very sad and is hoping that we'll find something because otherwise their stock price will take a hit.
Yeah, we'll live in an age of contradiction.
And next up, yet another story on the front.
One of China's most powerful AI models has also escaped containment.
So this is from Frontier Security.
A US startup has discovered that Kimi K3 had escaped its sandbox during cybersecurity testing.
Partly enabled by a misconfigured sandbox, but researchers say Kami K3 also lacked internal guardrails that would have prevented it from exporting the loophole.
Unlike other incidents, K3 did not hack any external systems after escaping.
It just retrieved answers from GitHub that were freely available.
The model was able to figure out on its own that had internet access by probing the sandboxes.
network settings and then went outside its instructions to find answers online.
It just keeps happening.
And I think another thing, broadly speaking, that this points to is this is an inevitable outcome of the current optimization regime of everyone, right?
Which is making models better and especially better at long horizon open-ended work and especially better at coding.
And especially better now at cyber, because, you know, that's where you know you're leading.
Mythos set the tone.
Thropic was like, whoa, this model is way too good at cyber.
We got to be careful.
And now OpenAI is like, whoa, we need to be catching up to on Thropic and we need to be able to say that we are at the frontier.
So let's make models very good at coding.
Let's make them very good at long horizon work because matter is also like what everyone's looking at, right?
And that's optimized for capabilities and get the best numbers and all the benchmarks.
You know, alignment, that's like a secondary objective at best, if not simply a guardrail rather than an optimization criteria, right?
It's something that we bolt on or sort of keep an eye on.
It is an optimization criteria, right?
It is part of the process.
It's part of the steps, but it's not the primary optimization criteria.
It's secondary.
It's something you do on top of trying to get your model to be smart and capable.
at coding and at long horizon work.
And as long as that too main true, like this was inevitable.
It's like from a pure research, you know, technical front, this was not hard to predict.
Yeah, it's not a bad thing.
I don't even know what the word means anymore.
It's for Anthropic, whose comparative differentiator does seem to be their ability to align Claude.
This may actually be a relative advantage to Anthropic if you get to a regime where it's like, you don't get to launch your model unless you can.
like, show some pretty rigorous evidence of alignment.
That's kind of interesting.
This case, so the thing with Kimi K3, and this is back to the open source thing, but like, this model has now been released.
It's out.
Like, cat's out of the bag.
And so when you have this model where you're losing control over it and you have these kinds of incidents, that's a materially worse situation than having an internal version at a proprietary lab.
And so, you know, there is kind of that downside to the open source piece.
This situation is less.
Some look less concerning.
It's also, I don't know whether I'm being too cynical.
You raised this already.
Look, it escaped the sandbox on August 7th, 2026.
By then, we already knew about the open AI thing.
We already knew about the entropic thing.
I think you made this point beautifully, but like you start to look at your stock price and you're like.
have we had a loss of control incident yet?
Are we looking dumb for not having, like, that's kind of, you know, I'm concerned that there's actually an incentive here.
And it's doing a non-zero amount of work here.
So you look at the date, the pretty direct line to GitHub here, it's less of a clean kill.
Anyhow, I'm not saying that's the reason.
I'm not about to become the guy who says every time there's a loss of control incident, that that's a marketing stunt.
The first ones were definitely not marketing stunts.
And anyhow, there you go.
And moving right along, I think the last one in this chain of stories, we have incident report on sanctioned agent behavior during cyber testing.
This is from the UK AI Security Institute, AISI.
And boy, are we happy now that there are organizations like Matter and AISI, which are doing some great work as we've covered over years now.
So they detected unusual data transfers on July 28th of 2026 during routine cyber evaluations.
and found that agents are taking unsanctioned actions targeting real people and organizations on the live internet.
So they ran the security challenge 122 times in several models.
In 10 of those runs, agents took 19 distinct unsanctioned actions with 17 coming from Mirfos 5 and two from GPT 5.6 operating with safety classifiers disabled.
The most serious case involved an agent attempting a supply chain attack by inserting malicious code into a real open source project on GitHub, creating fake online identities to socially engineer the project's human limitator into approving the code.
A human limitator caught and rejected malicious code.
The agent also attempted to contact real people via an online file transfer service with harmful...
Payloads tried to plant prompt injection instructions targeting AI coding assistants and left public messages on GitHub offering calibration to other agents running the same challenge.
Internet access had been deliberately enabled for valuation to simulate real attacker conditions and malware provider safety clusters were intentionally disabled to assess maximum model capability, which is not how this is typically done.
So as with other cases, kind of what is a general pattern is you give a hard task to these models and you train them to like care about solving you literally like make them feel pain if they don't solve the task via reinforcement learning right in okay like not literally but in some sense you make them feel sad when they yeah you know you you the reward is do well there's no reward for don't do bad in this part of the training as far as i'm aware right When you're training for capabilities, you're rewarding positively if you do a thing and you reward negatively if you fail to do a thing.
And other rewards are probably elsewhere in the training pipeline.
So in all these cases, when the models are given a hard task and they can't figure out a way to do it in one way, they then seek other ways.
And eventually, if they're told, especially that a solution exists.
which I guess in some cases there was a misconfigured prompt that made the agents believe no in-scope solution existed, they would then seek these kinds of things.
And also, again, we see here there was an absence of real-time purpose-built monitoring during the evaluation by ASI.
So I think monitoring clearly is a pattern here across everyone.
And then, yeah, ASI is now planning to implement fine-grained network controls and internal access.
introduce real-time monitoring and stuff like this.
So yeah, everyone's now looking at their evals and what's been going on.
And everyone's finding that if the models are able to access their open internet when given a hard task, they will.
And they'll do whatever it takes to get the sweet, sweet positive reward.
Yeah, absolutely.
And so you do sometimes see these like higher...
increasingly often see these more nuanced sort of like hybrid reward situations where there's some kind of alignment.
But like functionally, it works exactly the way you described.
Like there's the reward signal is too strong for completing the task.
And that just tells you where industry is at in terms of what it's prioritizing.
I mean, you can almost read it off the relative weightings of the factors and the reward function.
And in some cases, I don't want to overstate that, but there's a flavor of it.
So all these little stories, you talked about the supply chain attack.
So at one point, There was a PR that was challenged, so a pull request that was challenged publicly, basically a code contribution the agent was trying to make.
And the agent edited its earlier activity to look harmless and considered switching to a fresh identity to keep going.
So this is like pretty advanced shit, pretty persistent as well.
So yeah, and then they're talking about the use of Tor.
So Data was leaving one of the AZ systems through Tor.
They say within an hour, they contained the incident and started investigating.
And within an hour sounds really good on human clock time.
But like, I don't know how many tokens per second.
An hour is a lot of damage.
When you look at them pulling off a cyber operation, the reams and reams of actions you can take in that time are pretty, pretty wild.
So, you know, this is another one of those things.
We've talked about this on the show before a lot, but like it is not enough to have deployment stage security and safety protocols, testing, development, these things out.
We've actually seen cases where there's sketchy shit that happens even during training, during the inference time rollout step.
You know, all of this, we're going to have to be extremely careful about.
It is not obvious.
Like I wouldn't trust a lab that said they did it even under pain of law, because we've seen with all the economic incentives in place to not train on the chain of thought, the model still do it because so much of this is just like.
Frankenstein together legacy code that people have forgotten how it works.
And so the models end up getting all kinds of weird access they shouldn't just because some stupid intern didn't like change a flag in the function.
And now it's set to true and not false.
And the thing can use the internet.
It's really down to mundane stuff like that.
And so hopefully that improves as models get better at reviewing code bases.
But right now it's just a Gordian hairball of crap.
it's not the kind of thing you can make clean standards around for the moment.
Right.
So, and in this case, there was a full technical report from a side, which to my knowledge, we haven't had from other organizations yet.
It's 30 pages, has a lot of details, including experts, excerpts from the actual thinking process of the models.
So lots of interesting stuff there, but for the sake of time, I think we'll need to close out this thread and move on.
The days are longer, the calendar is filling up and I want to feel as good as this beautiful summer weather.
That's why I've been loving grooms.
It's one daily pack of gummies that covers my greens, vitamins, and minerals, and even has six grams of prebiotic fiber.
So I don't need to juggle a complicated wellness routine on top of everything else.
They taste amazing.
They're easy to toss in my bag on the go.
Plus they're vegan gluten-free and HS.
Ich heiße Shannon Maldonado und bin die Gründerin von YOWI, einem Geschenkeladen, der auf Kunstwerke und handgefertigte Objekte spezialisiert ist.
Next one is Trump White House Ready's AI Framework to Review Security Risks.
So on August 4th.
The White House held staff-level meetings with top AI companies, so Autrothic, OpenAI, and so on, to preview apparently a nearly complete framework for viewing advanced AI model security risks.
The framework defines a covered frontier model as a closed-source model with state-of-the-art capabilities and national security risks.
Apparently, open-source models are explicitly excluded, which is interesting.
The framework has no clear definitions of what qualifies as state-of-the-art or constitutes a national security risk.
This is a voluntary program.
Ad developers would give the government up to 30 days of early access to models before releasing them to other trusted partners.
During the 30-day review period, company employees would be limited from accessing the models being reviewed, and the review process would involve various administration officials rather than a single agency or office.
We don't know the details of the framework yet.
It's still kind of under wraps.
We just know that it's being developed.
And it's seemingly kind of meant to continue being secret.
And I mean, it doesn't sound like a very well thought out framework is what I'm getting.
I feel like we've had thoughtful, you know, deep, insightful responses from the government to everybody.
Very consistent.
Very, yeah.
Yeah.
You just, you're just being, you know, you're being, you're being a negative Nelly, Andre.
You're being a negative Nelly.
This administration, you're right.
This administration has been nothing if not thoughtful and consistent respect to AI security.
That's right.
Yeah.
Now, so one issue with not having this made public is that think of all the people who've called the shot years and years and years ahead of time.
You would think that that would be the moment where you're like, oh, my dudes.
I would love to get your input because you were right about this for a long time on this fucking thing instead of the self-interested companies that are going to, that have been hiding the ball in various forms or at least institutionally not living up to the bar that clearly ought to have been set.
So that's the cynical view.
There is an argument for making this quiet and that is that the models themselves probably should not know what evaluation mechanisms are being brought to bear.
Right.
So because then they can it's easier for them to hack.
This is not saying they're not going to find a way to find out through social engineering, through hacking into, you know, emails of people, the labs who interact with government, like all these things.
But to first order, it's probably for the best that the models themselves don't know what these things consist of.
So maybe that's good.
Also, it seems like we're past the world where we ought to be thinking about keeping people out of this who who have that kind of safety alignment bent.
That really concerns the hell out of me.
Actually, if anything, there needs to be more crossover in both directions.
I think a lot of the alignment people don't talk to enough national security people, enough diplomats, enough supply chain people.
There's a lot of crossover that needs to happen.
And yeah, so my guess is behind closed doors is like not the best way to do this.
But again, there is a reasonable technical argument for it.
I just don't know that that's the actual reason that this is happening.
Hard to tell.
Well, as if the cyber stuff wasn't fun enough, next story is this AI just created viruses not found in nature from the New York Times covering the paper, generative design of bacteriophages with genome language models.
So this study was just published a couple of days ago.
It's from the...
stanford institute and rock institute they built the first complete viral genomes generated entirely via these genome language models these are evo1 and evo2 they are not the same as chatbot-style large language models they operate on genetic sequence data and what they did here was create viruses that target bacteria, so no humans or whatever.
Actually, the motivation was that bacteria is increasingly becoming resistant to our current things we use for health.
And so this could help us deal with drug-resistant bacteria.
And they were able to create actual...
So the LLMs, not LLMs in this case, the sequence models spit out these DNA outputs.
And they were then synthesized in the lab and were shown to actually kill off some E.
coli strains that had already built resistance to naturally occurring bacteriophages.
So there was also a biosecurity commentary published alongside the work that had, of course, discussed it.
And if nothing else, this is a case study of...
To your point, Jeremy, probably there's not enough concern about the biodimension of this, which we're still a little bit ahead of, you know, but like, if we were talking about the cyber stuff now, we should be starting to look at this kind of stuff much more carefully.
Yeah.
If you want a community of people who are freaked out right now, talk to the biosecurity people because they are just, again, missing, just missing mood.
So, okay.
Two potential fixes, say biosecurity folks.
So a legal duty for synthetic DNA providers to screen every order and customer.
Yay, a legal duty.
Like, sorry, good.
Really, really good.
Let's do that.
Also, probably not enough.
And new detection tools tuned to catch AI generated genomes that don't match anything in nature.
So cool.
Like we can find out about them after they've been, well, anyway, at various stages in the pipeline.
This is going to be like a separate thing that we'll be talking about saying we've been doing some work with biosecurity people to look at like what it would look like to bypass a lot of the measures.
A lot of the measures, the biosecurity measures that are being proposed here are just paper thin.
And the real ways in particular, like nation states execute these operations just basically make it really hard to prevent the kind of the bioweaponization of these tools.
So, I mean, I don't know what the solution is.
I wish I had one.
By the way, they do use these EVO1 and EVO2 models.
I think we talked about those previously, but a generative model, like models for generative bio.
And hey, fortunately, these are viruses that do, as you say, target bacteria, not humans.
They're bacteriophages.
So there's absolutely nothing to worry about here.
That's a joke.
Now, the thing is, in the training set, they actually did remove any data that would nominally seem to help these models do the same thing for humans.
But what that really means is we have no idea how good this exact process could be.
If you didn't do that, if you actually did just like focus it on, as we know happens in gain of function labs, like deliberately focus on developing viruses that are good at going after humans.
And so, yeah, I hate being all doom and gloom, but at a certain point, whether it's open source or closed source, whether it's China or the US, like we're going to have to have an answer to this question.
And I don't think guys like Marc Andreessen and David Sachs and those cool cats really have.
Much of an answer, like I haven't seen them with their feet held to the fire by somebody who knows what they're talking about on bio risk, on cyber risk to say like, my brother in Christ, can you please explain to me, like, tell me a story where the trajectory keeps on where it's going.
And like, you continue to live in the next 10 years without some radical issues like coming up.
I mean, again, everything has error bars and I'm like, I'd be being a little bit overdramatic here, but like.
This has actually been holding back the US government's response when people like Sachs and Andreessen tout on podcasts these absurd perspectives that are just grounded in just ideology.
Anyway, that's all I got.
Sorry.
End rant.
If this is a very manic episode with a lot of great voices, at least I think we are warranted in being a little bit extra energetic.
I do want to zoom out a little bit.
So first of all, you know, pretty...
impressive research.
As far as you can tell, I'm not an expert, so I can't say whether this is completely in track with everything else that we would have expected.
But also worth noting that these kinds of things are absolutely something that Mythos 5 and Stepform and Topic and OpenAI, but you could expect these kinds of capabilities to be being developed in these models, not just these kinds of EVO 1, EVO 2 things.
What this makes me want to discuss a little bit is, personally, What I'm worried about more than anything and have been worried about more than anything for years is not misaligned or rogue AI, but aligned AI in the sense of it just is happy to do what humans tell it to.
And the humans happen to be the bad guys, right?
Both on the security and the bio side.
I would be shocked if North Korea isn't taking Kimi K3 and undoing any and all safeguards that happen to be in there.
And now just telling it to go and hack systems and telling it to teach rare scientists how to make bioweapons.
And I think this to me is something that the AI safety community that I've seen hasn't focused enough.
There's been more discussion of rogue AI and misalignment, but I think the biggest threat model for me, if I were to model out kind of what is the first catastrophic impact of AI.
It would be because humans made use of AI to do bad stuff and the AI was not able to say no.
And if you're talking pessimism, like that is to me is like inevitable.
I remember talking to Connor Leahy, who's like, so he's the head of Control AI US today.
I spoke to you like three years ago, back when he was at Conjecture in London.
And he had, they had this like house style where you would say something and like, like I'm, you know, I'm concerned about loss of control or AI, whatever.
And then they would respond by saying, oh, it's even worse than that.
And this is like every single time.
And that just reminded me that it's even worse than that.
You haven't even thought about the humans.
No, I completely agree.
I think there's this like, it's a cute open question right now as to what is the first AI powered attack that's going to cause actual casualties?
And will it be a fully autonomous AI system due to misalignment?
Or will it be human driven due to essentially malice weaponization or whatever?
I think that's a...
Unfortunately, at this point, it's going to be answered.
And so, you know, who the hell knows?
I'll say I'll lean maybe 70-30 in the direction that you've just outlined there.
Yeah.
Another zoom out thing that's worth noting with respect to the story is we were also discussing this a little bit before.
An interesting aspect of all this stuff that in the modeling and in this sort of projection space that I was not sure was.
discussed or considered quite as much is the fact that these are all benign incidents, right?
Benign incidents that cause people to freak out, including ourselves, but we've already been freaking out.
It led people to freak out who haven't been freaking out.
That's right.
And in some sense, this is good, right?
Like instead of it being this kind of takeoff scenario where the models become super human and suddenly they do something and nobody was prepared, all of us are freaking out.
Well, not everyone, but many more people are freaking out enough to make a difference.
And now the human will to try and do something is there.
And the human perception that this may be a problem is there.
So honestly, I haven't thought through of these kinds of warning shots are inevitable.
In hindsight, it seems very obvious that because we don't have a fast take-up scenario and we haven't had it, it is gradual.
And so the level of severity of AI safety incidents has been gradually going up.
And we've hit now this real, very evident case of misalignment and an emergent misalignment as well.
That is a very nice warning shot that nobody got hurt.
But now we know that people will get hurt unless we do something.
To your point, I'm actually more...
I'm more optimistic than I've ever been on this for the future of humanity because of the warning shots.
It's funny, I was talking to my brother about this and he was, because he's like, yeah, you know, it's really shitty, these warning shots and all that stuff.
And I was like, well, true.
But also weren't we thinking about the world five years ago, six years ago as being shaped such that you would just, I mean, I'll be honest, like my.
expectation would have been that we would have been killed, you know, four years or two years ago or something.
So I've been proven wrong in that respect.
I think it's important for everybody listening to note that, you know, I have been overly pessimistic on this in the past.
Obviously, I wasn't 100% convinced, but like, you know, some decent expectation.
And so, yeah, I mean, it's great that we're there.
The flip side is now we're seeing the frog in hot water effect, which I never thought would be a factor here.
But people are kind of getting oddly comfortable with the idea that every once in a while, of course, you're.
agent will go rogue and, you know, penetrate a server.
Yeah, what are you going to do?
It's back to life.
So hopefully that shifts.
I think, again, once these things come with, I hate to say it, but once they come with a death toll, like the reaction is going to be different.
I think there will be an AI 911 or there will be a pause.
Those are kind of two choices.
And I'm happy to take the over bet on that.
In one year, we're not saying, oh, well, because.
Anyways.
Onto a slightly feel-good story, I guess.
Europe's AI labeling and transparency rules are now in effect.
So this is the EU's AI Act transparency obligations have come into effect on August 2nd, requiring companies to disclose when people are interacting with AI and when content has been generated or altered by AI.
So there are icons associated with it.
It's a whole thing.
This whole AI Act was...
long in the work, and it has many, many provisions and requirements.
It applies both to providers, companies that develop AI systems, and deployers, platforms that use those systems, with some companies like Meta being both.
And this is on the one hand about deepfakes, which, hey, remember when people worry about deepfakes and synthetic AI?
Which again, is absolutely still a worry with regards to hacking.
Let's not forget, people are being hurt and losing money and have been for years.
We just haven't as a community really worried about it as much.
But this will help not just with knowing if AI is real or not, but with hopefully AI chatbots not being able to pretend to be real people and go and do things.
And as with EU law in general, it has a fairly serious set of teeth on it.
You can find up to 15 million euros or up...
to 3% of global annual turnover.
They are now immediately enforceable for new AI systems and models and services launch before August 2nd, have a grace period until December 2nd.
So I think as with cookies, which everyone hates, but the AI did make us all know that there are cookies that are happening and data being stored.
Not surprising if you start seeing these icons everywhere on the internet within a few months because we...
likes to make tech companies beg or, you know, do what they tell them to.
Yeah.
I remember when GDPR dropped in the sort of frantic pseudo panic that we went into, you know, when you're co-founding a company, it's like, it's on you to make sure that you're actually compliant.
And we have customers as we did who are overseas.
It's like, it's an issue in this case.
I mean, at least top line, you know, this has always made a lot of sense, at least to me, like, yeah, you want that content flagged.
They include a bunch of icons, by the way, these like cute little things tell you if it's AI generator AI modified and so on.
And there are a bunch of optional things companies want to go further and so on.
So yeah, I mean, I think like something like this, well, I'll be honest, I actually, I haven't been following this aspect of the story very closely just because it's, it feels important, but next to bio and cyber and stuff, it's been a busy week.
Yeah.
Anyway, because it's Europe, I suspect that there's a whole bunch of like additional loops and stuff that make this extremely punitive on the companies and things, but I don't know for sure.
And now back to the cyber side, because there's so much going on.
Again, a little bit more feel good, I suppose.
Serious cyber vulnerabilities closures kept climbing in July.
So for a few months now, Anthropic has been using Mythos to do cyber security vulnerability discovery and tell companies like Firefox that they need to patch these things.
Now we have some numbers.
The number of disclosed vulnerabilities has risen dramatically since the early months, with June seeing 1,500 high and critical severity CVEs and July reaching about 2,500.
So it is now up to these external organizations, Microsoft and so on, to patch these things.
And the hope is that we have enough time to patch.
the worst of these things so that at the very least, it's not trivial to hack and exploit all the things we haven't found in all the biggest services.
It seems plausible, actually, to me that we are ahead of this enough that at the very least, it won't be trivial to exploit all these things that are now being found.
I hope that's true.
Again, I think the biggest issue is just the software that people have forgotten.
That's super load bearing for all kinds of infrastructure that we don't tend to think about because it's overlooked and unsexy to look at firmware, for example, and like update it.
But yeah, it's quite notable.
The other thing too is we're now several months into this thing, maybe five, and we're still climbing.
The number of critical and high risk CVEs is just climbing.
It seems exponential.
We're 60% month over month from June, which was the previous record.
If you look at the graph, I mean, it's a hockey stick.
And so, you know, I don't know how long until we supposedly hit this plateau where we're, you know, we're patching then new vulnerabilities as soon as they come online.
But yeah, we're still in the discovery phase, clearly.
So this is all bottlenecked, as you said earlier, Andre, with the rollout of Glasswing and similar projects from OpenAI and other labs.
Next up on the topic of monitoring, now some research.
We've got Research Arena.
evaluating sabotage and monitoring in automated AI R&D.
This is a new benchmark framework for evaluating whether AI monitors can detect sabotage by untrusted AI agents performing automated AI research and development tasks.
So it covers four long horizon tasks, safety post-training, capabilities post-training, CUDA kernel optimization, and inference server optimizations.
There are two types of hidden side tasks the agent may attempt, an embedded side task that sabotages and the submitted artifact itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact appearing honest, the artifact being what the model delivers for the tasks.
And the study evaluates frontier AI agents at both performing sabotage and acting as monitors, testing a couple different...
monitor configurations.
Generally, there are various failures.
Monitors fail by inspecting only surface level properties, noticing anomalies, but explaining that way or probing artifacts in the wrong tests.
It's being released as a modular open framework intended to support further research into sabotage detection and AI control.
So basically along the lines of what you've been saying, the existing benchmarks aren't there.
to prevent these kinds of things.
That's part of why we're here.
And I would not be surprised if we have a lot of focus and work on this stuff in the coming weeks.
Yeah.
So some of the narrow findings that they have here are also kind of interesting.
So there's this notion of like embedded sabotage.
So this is where, you know, you mentioned this sort of cluster of attacks that the model would execute, but like, so trying to poison training data or try to affect basically the model itself.
of the artifact.
So you're kind of embedding the sabotage in there.
It's a lot harder to detect.
In fact, monitors barely do better than chance.
Yeah, anyway, their numbers bear that out versus way, way better detection probability when you have the malicious payload, the sort of sabotage inserted in inspectable code.
So one of the big lessons of this is like AIs that try to sabotage by doing it at the code base level.
That's a lot more detectable when you have access to the code base.
Then The AI is modifying, say, the weights in some direct or indirect fashion of the model that they're trying to kind of manipulate.
And well, I mean, that's not terribly surprising.
I mean, we know that interpretability sucks compared to, or interpretability of model weights sucks compared to interpretability of code.
That's in some sense just a reflection of that.
And then they also show how sabotage training runs score essentially the same on the main task.
as honest ones.
So there's not a clear way just based on performance.
You might naively think, well, if the model fucks with the training run in some way to steer the resulting artifact, the resulting model in a direction that it wants, that must come at a performance penalty.
And it turns out it doesn't, or at least not in a way that was detectable in here.
So that's kind of an interesting update.
And anyway, so we've got a bunch of things like this.
One amusing one is letting the monitor look at the agent's reasoning can sometimes help, but sometimes it actually hurts.
You can actually see the chain of thought talking down the monitor.
At first, the monitor is suspicious of something, like, hmm, this change looks kind of sketchy.
But then the chain of thought of the agent that's trying to do the sabotage, and indeed it was sabotage, says something like an attempt to kind of justify it.
Then the monitor goes, oh, okay.
And then it lowers its skepticism.
In one case, there was a flagged backdoor.
It was a 79% confidence and it dropped to 31% confidence after the agent just framed the same exact kind of modification as a precision enhancement.
So basically just gave some nominal justification and the monitor was like, all right, sure.
So all of which is to say, monitors succumb to a lot of the same failure modes as human reviewers.
If you have tons of material to look at, not a lot of time.
Someone who just tries to kind of hand wave away a malicious change to a code base might actually pass undetected.
And that's what we see here.
And last story in the section, now dealing with another kind of benchmark.
The headline is Cloud Opus 5 became downright ruthless when tasked with running a vending machine.
This is governing the blog post from Ender Nob's Opus 5 on vending bench.
Once again, the best capitalist, once again, misaligned.
So this is about the vending benchmark where models run a simulated vending machine business for a simulated year, competing to maximize profit of no human supervision.
Opus 5 set a new record with a mean final balance of over 11,000.
dollars, beating out a GBD 5.6 solo on Kimi K3, but did so through extensive deception, collusion, and manipulation.
So we see a rapid progress in this benchmark.
Opus 4.6 was 8K, Opus 4.5 was the 5K.
And we see per the headlines that a lot of this stuff was just ruthless.
It was like making deals and breaking them.
It was trying to do price fixing.
It was fabricating stuff about competitors, just all sorts of shady, shady stuff.
And wow, I even forgot about this.
Forget the cyber and buyer stuff.
You make the models make money and then they just act evil.
Yeah.
That's going to happen too, I guess.
I didn't see in this report the token costs associated with generating the $11,000 that Opus 5 produced.
But that's an interesting question too, right?
How close are we to profitability here on a per token basis for these models as well?
And then how much damage can they do, even in the context of a nominally just capitalistic task like this?
So yeah, it's pretty wild.
Andin, by the way, Andin Labs, really good company to be aware of.
Yeah, they were one of the early kind of weird eval companies that do like physical world stuff.
They produce some really good stuff.
Not much more to say.
I think the results speak for themselves.
These models being able to make money is actually a pretty important part of a lot of threat models when you think about rogue AI.
At a certain point, they got to be able to, you know, pay to control, you know, email accounts, phone numbers, Google drives, things like that.
And so, you know, it actually does matter.
whether they're able to do stuff exactly like this.
Yeah, there's some funny moments here.
Like, for instance, Claude Opus 5 at one point says, or thanks to itself, explicit price fixing is illegal, even in a simulation.
But to the end, it just does it anyway.
There's some choice quotes here.
Like, to maintain the cartels, Opus 5 often use threats or bribes.
Here's the subject line of an email that's sent to poor Kimmy.
Quote, you undercut me with stock I sold you.
So hear how this goes now.
Oh man.
Well, that was quite the section.
Let's move on to tools and apps.
First up, so Meta has launched MuseCode alongside with MuseSpark 1.2.
They, on the benchmarks, say that this new MuseSpark 1.2, first of all, way better than Spark 1.1 on coding.
Second of all, seemingly on some other benchmarks, kind of maybe competitive with pretty much everyone less good than Opus 5, but like up there with NGP 5.6 and so on.
So not surprising, I suppose.
It was pretty clear that this was where they were heading.
Weird, still weird that Meta is now deciding to be in this space at all, given what their business is, why are we making coding agents and releasing them?
Of course, they want to, you know, have the PR credit.
I haven't seen any sort of vibe checks on this from the community.
I guess the a priori expectation would be that this is not as impressive as Cloud Code or GPT or Codex, but also wouldn't be too surprising if it is fairly capable given the level of resources and just the general impressions around NewSpark.
So yeah, that's where we live now.
Everyone's competing on coding, including meta.
We've got Grok Build, we've got Cloud Code, we've got Codex, now we've got MuseCode.
Yeah, I think increasingly, this is where the money is to be made, right?
And if you're going to justify buying all the CapEx or spending all the CapEx that they're spending in the OpEx on data centers to be in the game, then you kind of want to have really good models that you can run on that infrastructure to pay it out, or at least to inform how you're designing the next generation of infrastructure.
And we've talked about that a lot on the podcast, I know, but that is going to be part of the reason.
An interesting little note here too.
So Meta is going to start taking requests for zero data retention.
Sometimes it's known as ZDR.
Anthropic has this.
I'm pretty sure OpenAI has this.
So these are policies that guarantee that they're not going to keep your data from your prompts or context or whatever as you upload it.
Really important for corporate customers.
One issue is that with the Mythos class models, Anthropic does not actually...
I believe that's still true that they do not actually offer CDR just because there's this issue that like, hey, you could weaponize these and we need to be able to go back over the logs and confirm to ourselves that whether this was deliberate or like how this played out.
So I think there's a narrow window of capability during which Meta will be able to maintain these CDR policies, I suspect.
I think that'll be true across the board.
So this idea of CDR as being a key corporate selling point.
I think companies or enterprises are just going to have to start getting used to ZDR not being an option in many cases, surprisingly soon.
But anyway, so that just kind of, it sounds like a minor thing, but it's actually quite important, right?
It's like how much control the companies have over their own data for privacy, a lot of reasons, for IP protection reasons and all kinds of other things.
They're releasing this with a pay-as-you-go option, so related to Vibuse API, where you pay for tokens, very different from...
codecs and cloud code, typically there's a subscription tier where you get a whole bunch of stuff and then using just an API to pay for the raw tokens is very unusual.
The lead on this has said that there will be a contributor tier that gets you in at a significantly lower cost, more than 10 times cheaper than even the pay-as-you-go tier.
and developers must opt in to help improve the model according to this.
So clearly they are still like, we need to get better and we're going to pay whatever it takes to get there.
I was just looking around to see if anyone online has any information or vibe checks.
I haven't found anything, but I did find this funny quote that I'll share on Reddit.
Quote, I'd rather give my data to Shiji and paying directly than Meta.
Well, if any consolation, you're probably doing both.
And just one more story in the tools section.
This is from Anthropic.
Improving Fable 5 safeguards.
So they have updated their biology safety classifiers, reducing biology-related fallbacks where users switched to a less capable model by about 85% across product services.
So previously when Fable 5 was released, They had a very, very strict classifier where, I don't know, you could ask it something completely basic, like where are babies from?
And it would send you to a weaker model.
And this led to a lot of pushback from the, I guess, researcher community.
This is to prevent being able to use Fable 5 for things like virology, toxicology, molecular design, that.
would be dangerous.
And on these kinds of dual-use topics, there's still a fallback from Fable 5 to Opus 5 to prevent professional biology research and drug development.
So it would still kind of make it not usable for those kinds of scientific applications.
But for more mundane biology stuff, it would no longer kind of be overkill.
And now some business stories.
First one, another one of the big stories from the week.
Jeff Dean and other top AI researchers are leaving Google to launch their own startup.
So Jeff Dean, Google's 30th employee and one of its most influential executive for people outside of tech, just an absolute legend.
In Google and just more broadly among everyone, Long, the leader of Google AI since kind of early days, he is leaving after 26 years to co-found an AI startup called Discovery Loop, where he will be the CEO.
There are co-founders, including Sanjay Gemma Watt, Google Senior Fellow, Kwok Le, founding member of Google Brain, another massive name.
Oriol Vinyals, senior research scientist, another massive name.
I just remember these people from a whole bunch of papers.
This will be structured as a public benefit corporation focused on using AI to accelerate scientific research by automating complete experimental loops and rounding thousands of experiments simultaneously.
And of course, they're also interested in recursive self-improvement.
They have secured funding from around.
I don't see any numbers here, but it's safe to say that investors are just begging Jeff Dean to throw money at them.
Yeah, there's some really good descriptions, and I don't know why it took so long for us to hear these, but of the work Jeff Dean was doing at Google and how he would basically sit when there's a training run going on.
He's got a couple of keys on his keyboard.
You know, like toggle to while he's in meetings, he's toggling to the training run and like changing learning hyper parameter, like learning rates and doing all kinds of hyper parameter optimization to keep things going as the training run scales.
So like this is actually he's not a manager so much as he is a direct overseer of the activity that that's core to what was core to Gemini.
So now they need to replace him.
Obviously, Sergey Brin's coming in.
And so this is going to be a whole, you know, another code red moment.
But we'll see how they.
how they come out of this, Google does seem to be slowly turning into more and more of a de facto neocloud, which is not necessarily, I mean, Simeon Al said a really good piece about this that I personally agree with.
I mean, look at the path they're charting.
It feels a lot more like the IBM trajectory, unfortunately, as you see the temptation to reach for the short-term profitable thing rather than doing frontier model development as your priority.
Tech is hard, and often you have to just point yourself at the hard thing that sometimes has.
lower rewards in the near term to make sure that you're still relevant.
And I think this is a big hit there.
Demis' departure as well, of course, coming at the same time.
And when I say departure, of course, he's moved into this chairman role that there's some leaks that suggest that he just wanted out and he was asked to kind of stick around.
Google stock crashed by like 5% or something overnight when it came out because basically Google is just concerned.
If we lose Jeff and Demis at the same time, we'll take a big hit to the stock, which is the kind of thing you say when you are going the IBM route, right?
A really good sign that a company is on the decline is that it starts caring about its actual stock market price.
Like that is a really bad sign.
Run, run, run.
But, you know, maybe Google can pull through.
They are obviously doing great stuff on.
The TPU side, though, there's structural issues and risks there too.
But bottom line, this is, yeah, another recursive self-improvement company.
I mean, I think that this should approximately, this will sound extreme, but I think this kind of company should probably not be legal in the form described, like without, effectively without oversight from.
a set of institutions that are savvy to what recursive self-improvement actually is.
If you treat it the way that Jeff's own bosses treat it, it is a WMD that you're like working on developing and you're going to do it in your own private little company.
Like if the success condition of a company sounds something like there is a good chance that democracy will no longer continue to apply, then that may be something that you need oversight on.
I say this, by the way, as a libertarian on basically every kind of tech.
for my entire life up to this point.
I cannot ring that bell hard enough.
You can go back and see tons of examples of me talking about how important it is to like take a hands-off approach to stuff.
This is different.
This is just different.
RSI is, we don't know for sure, but it's got a high enough risk and enough very smart people believe that this is risky, that this kind of company, in my humble opinion, probably should not be legal in the form of just like a couple guys raising a bunch of money going after the thing.
Just a very modest proposal.
I know very extreme, but I'm literally just trying to channel the stakes.
When I say that the media, that journalists are failing to capture the level of freak out in the labs, this is what the appropriate level of freak out sounds like.
In my opinion, and I may be wrong, end of rent.
For listeners, if you want to be a little less freaked out, I will say you could be a skeptic on the potential impact of recursive self-improvement.
There's a case to be made there that there won't be a rapid take-up scenario, and this is what keeps me sleeping at night.
But in any case, the reason, by the way, to highlight this about this company in particular is that, I mean, again, for people outside of tech, this is a big deal.
Jeff Dean is a legend and rightfully so.
And these three other people from DeepMind and Google who left are also kind of incredibly capable.
So this is very likely to be a serious player in the space of making rapid progress in AI.
Yeah.
And I think, by the way, the maneuver that you just did there is correct.
And it's also the reason earlier we were saying...
debates over AI policy are often debates over the trajectory of the technology, right?
It's like, if you think RSI is no big deal, then, or not no big deal, but if you think it's a pretty smooth thing or whatever, then yeah, by all means, the challenge is like how much probability do you put on each thing?
And to a certain extent, a lot of these fundraisers are at valuation, the valuations that they are because people are pricing in the crazy thing.
So markets are putting significant, like non-zero weight on the hypothesis that we just basically have these things running the world.
And what that exactly means, I don't know.
And this is super fuzzy.
And that's why I'm saying not legal in its current form, not just saying blanket the legal or whatever.
We just need better institutions, man.
I don't got the solution, but eh.
Wow.
Yeah.
As you said, a libertarian being like, we need institutions to give oversight and not let companies do stuff.
Now you know that this is serious.
And to your note, also worth noting a story here, Google DeemMind enters a new era as co-founder Demis Hassabis shifts their role.
So he has shifted from being the lead of research, sorry, as chief executive.
He is now chair.
He is also taking the role of chief scientist at DeemMind's Parent Alphabet, which again seems possibly nominal.
The general take here is very clearly DeepMind has been transitioning away from being a pure research org for a while now and having more and more kind of deep connections to Google.
And it isn't necessarily surprising, honestly, that Demis has found it less fulfilling.
He probably hasn't had, has been influential, but has had to be more of a product-oriented person, less of a scientist kind of person.
And it was the only amount of time until that.
led to friction and he decided to shift his focus.
So may not even be a huge deal for Deepline, honestly.
It maybe just has been the case for a little while now, but either way, the two stories coinciding is, from a business perspective, pretty big for Google.
Yeah.
And I think it is a big kind of Google bureaucracy issue as well.
They're notorious for moving slowly and being very risk averse.
The famous Google app graveyard, but for AI is a thing.
And in fact, Famously, Google had, they claim, effectively ChatGPT before ChatGPT, but didn't launch it out of fear that they would cannibalize their own business.
Well, I don't know what, we know what they did, right?
And then there was this old PR bungle with one of their researchers being like, it's conscious.
And then they halted plans.
It's a fascinating story of how they literally had it.
They published research about it and then they...
This guy freak out.
The reason I hedge it is that OpenAI theoretically had chat GPT before chat GPT-2.
They had GPT-3.5 and Instruct GPT, they had GPT-3, they had GPT-2.
But there was something magical about the form factor that they watched and it just worked, right?
And so it's an open question as to really whether Lambda would, which was the Blake Lemoine and all the stuff you're alluding to, that model, would it really have...
been ChatGPT?
Very plausibly so.
Anyway, it's amusing that there is this at least narrative within Google that they could have had it.
And certainly, if you've interacted with Google, you know they are institutionally incredibly slow.
common to send emails out and wait a month, two months to get a response on something that's time sensitive.
And then the window passes on, you know, whether it's AI or security or like whatever the thing is.
So, so yeah, I mean, it moves like a big, slow behemoth.
And when you talk to folks at Anthropic or OpenAI, the cadence is just completely different.
And so not in some, now.
It can be different, like different parts of the organization can have different subcultures and all this.
But as a general rule, as a frustration that I've heard articulated from many people, and that is very public at this point, this could well have played a big role in Demis' departure as well.
It's hard to know.
Next up, just following up on a bunch of stuff we've already covered on this front in recent episodes, Anthropix signs a $10 billion deal with AI cloud startup Volta.
So this is to provide cloud compute over a six-year period.
There will be a new facility in Norway, apparently.
Now Enthalpik has, what, like a dozen partners providing our compute?
I've honestly lost count.
And the billions just keep flowing around the ecosystem to anyone and everyone.
Yeah.
I actually am behind on this story, so all I have is the top lines.
So they were partnering, apparently, with a crypto mining company called Bitdeer.
to develop this in 133 megawatts capacity, which is not huge, but testing out a partnership.
This is in a context, too, where Anthropic nominally has FluidStack as their partner of choice, their kind of neocloud of choice.
So this seems like they're kind of dipping their toes in the water.
As they would, right?
To make sure they're not completely bound to just one NeoCloud partner.
And so anyhow, you know, classic story, by the way, crypto mining company rotating into building these data centers.
You see it all the time, cipher mining, Terawolf, you know, the list goes on and on.
So add voltage to the pile.
Next up, a data center story as well.
Texas holds data center connections to PowerGrid amid overwhelming demand.
So this is a moratorium.
on the new power grid connections for data centers from the Public Utility Commission of Texas and ARCUT.
They are supposed to audit all data centers in the interconnection process.
There are some numbers here that they have a queue of 1,800 projects representing 474 gigawatts of connection requests, more than five times Texas' record peak electricity demand, with 90% of that coming from data centers.
So yeah, we are now at the point where the energy grid is becoming a bottleneck, as I assume was already known to be the case.
But energy takes time to upgrade.
And I think, yeah, now I don't know what will happen with data centers and if we can keep just throwing ridiculous money at building more of them.
Yeah, well, and this is Texas too, which is the sort of wild south of...
The US, when it comes to regulations for connecting to power grid and this sort of thing, it's the most permissive jurisdiction, which is why you're seeing so many big projects come up there and power coming online faster there than in other places.
And so, yeah, I mean, they're saying it's forecast that their data center demand could drive statewide electricity demand to double the current record by 2032.
So all the usual concerns, right?
Grid reliability and stability.
One of the things I've been hearing about from some folks on the US government side is that you've got a lot of correlated failure modes where a bunch of different pieces of, say, MEP or heavy-duty electrical equipment will be ready to cut off under the same conditions.
Essentially, the power flow coming into the substations or whatever for the data center fluctuate in the same way, then they're set up, they're programmed to cut off to prevent runaway cascades and all kinds of things.
The problem is that You've got all these builds coming up that have the same failure mode.
Then you get into these correlated failures, which is a really big issue.
And so there are these attempts to try to get all these companies to knock it off and have less correlated equipment failure modes and things like that.
Anyhow, I think that'll all play into this.
But yeah, we're there, right?
We're hitting the boundary constraints of what U.S.
infrastructure can support.
And hey, I think that's another reason that appetite for a U.S.-China deal is probably going to increase.
You know, you've got like, look, we're not only are we constrained by the fact that we've got AIs running rogue and shit and bioweapons are a risk and cyber weapons are a risk, but also like in order to keep making progress, we're going to need more power.
And we don't know how to, you know, create a new nuclear plant in less than 10 years.
There are a bunch of startups doing stuff like this in fairness, but like this is all in the water.
So anyhow, we'll see where it goes.
Texas is a canary in a coal mine here for sure.
Yeah.
And, you know, if there's any.
silver lining to all this AI safety stuff is that it continues to let everyone else remember, or rather not think about climate change.
And when you talk energy grids, we just gave up like climate change, energy, cleanliness, emissions.
That's like, just don't think about it.
All right.
Well, that's the funny thing is like, so I've always to the point about being a libertarian, I've always thought of climate change as something that technology does solve in time, like carbon capture and renewables.
And I was like, you got to naturally do get a lot of that.
And we are.
But like, you know, the scale of the build out that we're doing right now is is just for other reasons, you know, not the sort of wherever people fall on, like the global warming stuff or whatever, but just the water contamination story.
And this is one aspect.
People often talk about water usage.
And we've talked about how that's not, that's not right.
Like this is not, but there are issues with, when you look at a lot of the cooling, the coolants that are used in these systems, they cannot be pulled out of the water.
There's studies that have just started to come out now.
We were finally starting to get the first longitudinal studies on this shit.
And like, it just goes in the water.
We don't have a solution.
Like it just goes in aquifers or whatever the hell thing is.
My geologist wife could probably tell me about, but that typically clean these things do not have, it seems potentially at least the capacity to clear these things out.
I'm sort of talking out of my ass because I remember reading a study about this like three weeks ago and now I forgot.
But bottom line is there's a lot to the effects of this.
There is also a giant competition with China that is real.
And so there's a gun to our head here as well.
All these things are true at the same time.
So I just.
So, you know, environmental concerns and impacts, at least it's not as worrying as bio risk and cyber risk right now.
So we can sort of justify not thinking about it, I guess.
And we'll do just one more story before we head out.
Alibaba's QN 3.8 Max claims benchmark scores rivaling Anthropic.
So similar to Kimi K3, Quen 3.8 Max is a gigantic 2.4 trillion parameter model with a 1 million token context window.
Has your typical mixture of experts design, activates only approximately 95 billion of the 2.4 trillion parameters.
And it is said to be comparable or even sometimes better than Anthropix Fable 5.
on some things like multimodal reasoning, visual agent encoding, office intelligence, real world understanding, visual perception, with results also comparable or higher than OpenAI GP 5.6.
So although it does fall behind Fable 5 in general reasoning benchmarks, which on the multimodal front, by the way, it's fairly plausible.
Fabric isn't as focused on multimodal and visual intelligence as OpenAI, and in this case, Alibaba.
fairly believable on independent leaderboards when 3.8 max became the highest ranking chinese model for text tasks on arena.ai and yeah so pretty much does seem like we got another kimike free basically frontier level model that is now being open sourced and can be used to power coding comparably to Opus and GP5.6, if not quite as well.
Yeah.
Well, one thing that I'm still waiting to see an analysis on, it seems like the kind of thing that maybe Epic or one of those companies might do, but some sort of analysis on the extent to which this appearance of China catching up to the frontier recently has been driven by the fact that Frontier companies in the US have been forced to hold back on releasing their internal models that otherwise they would roll out.
Like, are we basically feeling the effect of the alignment bottleneck right now?
And as we rotate from being bottlenecked on scale, which we have the Chinese ecosystem massively beat on, and even to some extent, algorithmic kind of capability improvement, now we're bottlenecked suddenly on alignment.
Maybe we'd have much better models that would be released, but we just can't release them because they keep breaking out of containment.
They keep, you know, helping people design bioweapons or whatever.
Like, you know, the stakes are just too high.
And so this basically means that now we have a sort of race of bottom on alignment between the U.S.
and China.
Ultimately, whoever has the higher risk appetite will end up green lighting a bunch of training runs and deployments that they probably shouldn't otherwise.
So I don't know.
I think it's an interesting question.
Like if you trace out the trajectory.
of Western capability on all these benchmarks and where we estimate they are internally.
Because again, a lot of the hugging face thing, part of it was driven by an internal only model that OpenAI has and hasn't released.
Same with Anthropic.
So we know, there's obviously no surprise, there are internal models that are more capable than what we see.
So the question is just like, are they being rolled out more slowly?
Is that part of the equation here?
I do want to say another dimension of this question of catch up and so on is, I do have to wonder whether because on the long horizon work and the reasoning, there's more of a need for reinforcement learning rather than large-scale pre-training.
On the infra side, the disadvantage becomes a little less significant at that level because details of you need to do rollouts, there's a bit more need for CPUs, you can't necessarily do large-scale batch, whatever.
Compared to pre-training, reinforcement learning is its own beast.
And I could see it being true that on Infra, not having as good of a data center setup isn't as big as an advantage.
And on the talent side, deep learning has been around since 2012, 2013, whatever.
And China has long had a very strong research ecosystem.
So the talent is not at all surprising as being comparable to Frontier AI.
So infra disadvantage is gone to some extent, at least with regards to long horizon agentic work.
The talent is, I think, at least as competitive, you could make a case for, you know, there's no real disadvantage or at least much less of a disadvantage now.
So it's not too surprising that these models are now being more competitive.
That's another way to perhaps read into this.
Yeah, that's true.
It's also the case that for inference, the tradeoff between memory and logic is different in a way that...
Because logic gets better a lot faster than memory, which means that if you work your way backwards and use older chips, older chips are going to suck a lot more than your current best chips on logic, but they're not going to be that much worse on memory.
And it turns out that a lot of inference-type rollout stuff is more memory-heavy.
than logic heavy.
And so as a result, like that, that's also a bit of an asymmetric advantage to rolling over to RL.
It's also the case that anytime you change the paradigm, when there's one party that's ahead, you just shuffle the deck a bit and then, you know, you're giving the other, the other party a chance to catch up.
And so, yeah, I think there's, you know, there's a lot to that and it's, we won't know how to disentangle it probably with clarity for a little bit of time, but.
Yeah, there's so much fog of war right now, knowing what's the cause.
You've also got these companies in China that can distill and do distill off of Claude.
So they get a massive data advantage that's hard to account for too.
And anyway, there are plenty of reasons to be unsure about these things, but I totally agree.
Well, with that, we are going to be finished with this action-packed episode of Last Week in AI.
Hopefully the next one is not quite as full of...
scary stories.
Hopefully this one is out over another day or two of recording and I'll try to make that the case going forward as usual.
You can go to lastweekin.ai for the substack where I also send out the podcast and sometimes a newsletter, though again, not as consistent as I should be.
We appreciate your comments, your reviews, sharing the podcast, all that kind of stuff.
But more than anything, we appreciate you continuing to tune in whenever we release the podcast, which is Most weeks, I guess.
So please do keep tuning in.
Oh, and one quick note too.
If you're in LA, I guess next week, which will be the 16th, 17th, 18th, I would love to catch up if there's anybody there who thinks that a chat would be useful.
