# AI Market Shifts: Compute Commoditization, Open Source, and Regulatory Compliance

**Podcast:** Last Week in AI
**Published:** 2026-08-03

## Transcript

Hello and welcome to the Last Week in AI podcast where you can hear a chat about what's going on with AI.
As usual, in this episode, we will summarize and discuss some of last week's most interesting AI news.
Also, the week before, we have unfortunately skipped a week due to scheduling conflicts, but we will cover everything relevant from the period.
I'm one of your regular hosts, Andrey Kerenkov.
I studied AI in grad school and now work at the startup Astrocade.
And everybody, what's up?
My name is Jeremy.
Of course, I'm your other co-host.
I'm from Gladstone AI.
I do AI, national security, super intelligence, C-type, end of the world stuff.
So I sound a little thick right now, by the way, which is related to the reason that we didn't record the episode last week, which was that I was traveling and I got sick on the flight.
There was a guy who was coughing up a lung next to me.
And anyway, that's why I sound so weird right now.
But the trip was really useful.
And yeah, I'm going to hopefully be able to talk about a lot of this stuff soon.
A lot of conversations with like researchers at the Frontier Labs and folks on the safety teams, the capability teams, all that kind of thing that I think bears quite a bit on the events of last week and the week before.
We'll definitely be talking about a lot of that stuff with some of the inside view, a little bit what I can share right now on those things.
But man, things are moving.
And it has been slightly eventful two weeks.
I mean, I guess it's not the most eventful you've had this year, but there's been some big stuff that we'll be touching on as a quick preview.
There's a few new models, nothing gigantic, but fairly meaningful we'll start with.
Then, as usual, some funding stories and deals about compute and so on.
Some major open source releases, including Kimi K3, we haven't discussed about, so we'll be talking about that.
Then policy safety.
Of course, we'll be talking about the recent hacking incident from OpenAI and a whole bunch of other stuff related to that.
It's going to be a kind of policy safety heavy episode.
And then we'll round it out with some research and synthetic media and art.
So it'll be a packed episode.
This episode is brought to you by OutShift, Cisco's incubation engine.
Today's AI engines operate in silos, limiting their true potential.
We focus on building bigger, smarter models, but scaling up is just one approach.
To reach superintelligence together, we need to do more.
We need to scale out.
And we actually have a blueprint from 70,000 years ago.
Humans didn't just get smarter individually.
The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation.
Agents are now at the same inflection point.
They can connect, but they can't think together.
That's why OutShift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence.
By creating an open, interoperable infrastructure, OutShift is enabling agents and humans to share intent.
context, and reasoning.
The cognitive evolution for agents is here.
Explore internet of cognition at outshift.com.
That's outshift.com.
You'd like to thank Box for being a sponsor.
The key to unlocking the power of AI for your business isn't in an LM or an agent, it's the content stored in files across your company.
AI tools may be great at public knowledge, but they don't know your business, your product roadmap, sales materials, HR policies, and financial models.
And that's where Box comes in.
Box is building the intelligent content management platform for the AI era, serving as the secure, essential context layer for Box's AI agents to access the unique, institutional knowledge that makes a company run.
And that's a key idea.
The power of AI doesn't come from the model alone.
It comes from giving AI access to the right enterprise content.
Box's recent State of AI in the Enterprise Report found that 96% of organizations say agents need access to company-specific content, but only 36% have connected agents to trusted content across many use cases.
Box goes beyond file storage.
It connects content to people, apps, and AI agents so teams can turn information into action.
With tools like Box Agent, Box Extract, Box Hubs, and more, organizations can accelerate knowledge work, pull intelligence from unstructured content, and automate workflows.
And all of that is done with security, compliance, governance, and threat protection in mind, so employees and agents have only access to information they're authorized to use.
If you're thinking seriously about your company's AI transformation journey, think beyond the model.
Your business lives in your content.
Box helps you bring that content securely into the AI era.
Visit box.com slash AI to learn more.
And we'll go ahead and get into it, starting with tools and apps.
And here we begin with Anthropik releasing Cloud Opus 5, which they say comes close to the capabilities of Cloud Fable 5 in many domains.
and is cheaper, of course.
So this is following up on the release of Fable 5 a little while ago, Fable being their new family of models that they didn't have before.
And they also released Sonnet 5 either before or around the same time as Opus 5.
So they sort of caught up, presumably because these are distillations of Fable and Mythos, right?
So typically what you can predict with Anthropic is their big, kind of best model is their most compute-heavy, most impressive model, which is Mythos right now.
And these things like Fable, Opus, Sonnet are kind of derived from it to some extent, where they try to extract out the intelligence at a lower price.
So I think not a ton to say about this one beyond that it's supposedly quite good and close to Fable 5, so pretty big jumps in the benchmarks.
relative to Opus 4.8.
The vibe check has been a bit mixed as far as I've seen.
People have their usual sort of complaints about what the models are doing, and it's hard to know whether we just have high expectations now or, in fact, the models are getting stupid.
There's also a new fast mode in Research Preview, which offers you higher speeds at double the price.
Amen on, I think we're losing track of what it is.
that we're looking for just because the waterline is rising so fast, people are getting used to incredible levels of capability.
You're right.
This is probably a distillate of Table 5 or Mythos 5, then added safeties and all this stuff.
There's obviously additionally post-training that gets done to kind of further refine the character of the model after that.
And so one of the key things that they highlight here is the differentiator of Opus 5 is supposed to be more emphasis on verification and judgment.
So less kind of raw capability.
but more sort of like double checking its work, making sure that what you're getting is actually correct.
And so they give this example where like it's given a drawing of a machine part, but no way to view the original image.
And then Opus 5 like rewrites its own computer vision pipeline to extract the geometry from the raw pixels and reconstruct the part.
Basically the idea being like, it's going to get you that raw data, the original to base its conclusion on so that it knows it's right no matter what.
That's kind of like the vibe here.
Also in alignment and safety, kind of interesting.
This is always the game.
We live in a world where the U.S.
government decided to snap a chalk line at Mythos level and anything above Mythos level magically is subject to the de facto licensing regime that we have in the U.S.
And so in this case, Anthropic is in a hurry to say that their model is close to Mythos 5 at identifying software vulnerabilities, but it's less successful at developing exploit.
Again, that's part of the post-training.
That's part of making sure.
And also just as well, the pre-training or other parts of the training process where they're avoiding explicitly training it on cyber tasks in that way.
And so the goal here is really to position it as this is a really intelligent model that it is okay for us to release.
Of course, Fable 5 is okay, but it has additional safety over Mythos.
But, you know, you're always now going to see that kind of background concern of this threshold, which is, I mean, it's good.
It would be just great to have a more principled approach to the stuff.
One thing to note too, its performance on Frontier Bench.
So Frontier Bench is this new benchmark that we have.
It was released just a few days ago and the same team behind Terminal Bench came out with it.
It's this big community effort and it is basically just like a harder agentic environment.
We keep needing more and more difficult agentic evals to be like, all right, but you know, we just saturated your Terminal Bench or whatever.
Let's now move on to Terminal Bench 2, Terminal Bench 3 and ultimately this and there you go.
So here.
We do know that this particular model, Opus 5, is outperforming all other models on a cost per task basis.
And that's really where they're trying to differentiate it is cost per task.
Not necessarily frontier of intelligence, it's near the frontier, but cheaper per unit of intelligence, let's say.
At that point, it is slightly cheaper than GPT-5.6 Sol, the biggest and best model from OpenAI.
And I think maybe indicative of like...
Pricing becoming more of a concern for customers at the business front.
Now that there is a competitor to Anthropic with Codex and OpenAI being quite capable and very, very cheap alternatives for model usage, I have to wonder whether we're going to be seeing more kind of pricing pressure going on.
Next up, some more model releases, this time from Google.
Google DeepMind has released three new AI models, Gemini 3.6 Flash, Gemini 3.5 Flash Lite, and Gemini 3.5 Flash Cyber.
So per the Flash aspect, these are cheaper and faster than, let's say, more intelligent models.
price.
Gemini 3.6 Flash is priced $1.5 per million input tokens and 7.5 per million output tokens.
So that's slightly cheaper than Sonnet 5 and not super cheap.
And then there is, of course, the cybersecurity-focused Gemini 3.5 Flash Cyber, which is integrated into this code vendor agent that autonomously builds exploit code to verify vulnerabilities in sandbox environments.
and then generate patches, which they say has found problems in complex real pieces of software, such as the V8 JavaScript engine.
So I think it's interesting to see a general movement towards cyber focus with not just the Mythos and Opus and so on.
We've seen OpenAI release a cyber model.
Now Google has released a cyber model, and we'll be discussing Microsoft has also released.
a cyber model.
So I think everyone is like, oh no, we got to do something about this.
I don't think it's sharing too much to say like people in the national security space and the frontier labs are really concerned about where cyber is going.
And this view that, you know, we're in the vulnpocalypse right now, right?
We're getting all these low-hanging fruit vulnerabilities being discovered and exploited.
Assume we're going to get a mythos class open source model sometime in the next, you know, certainly six months, maybe a bit less.
At that point, you're going to need an answer.
And so all the labs are pre-positioning for the moment when basically they're holding the world for ransom.
I mean, you know, you need to use really, really good cyber models to shore up your infrastructure or else.
Like that is just going to be the case.
As an aside, I'll just like casually drop the prediction here that we may see some pretty significant disruptive cyber attacks at massive scale, not even just nation state or their proxies, but literally just like disaffected young people or terrorist groups or whatever.
That's just what happens when you open source that level of capability.
It's just like, that's what the math says.
We'll see where that goes.
But that's like the default assumption right now of a lot of the people in the space, both on national security and the frontier lab side.
So that's part of what the positioning is here.
If you don't have an answer to the cyber question, you know, you're going to be a lot less relevant in the next six months or so.
They do.
This is pretty impressive.
I mean, the flash cyber model, which is maybe the one, at least it's the one I'm paying most attention to, is performing on par with a lot of.
frontier agents at things like, you know, CyberGem, and that's an important benchmark.
I mean, it's a lot cheaper too, right?
Really a fraction of the cost.
And so this idea of how cyber plays out is always a really strong function of how much compute you have.
You're a defender.
You have a certain pile of test-time compute.
The attacker has a certain pile of test-time compute.
Can you invest more test-time compute than an attacker to shore up your infrastructure is the question.
I mean, you know, there are a lot of ways to answer it.
you know, different ways to use test time compute and a question about how much leverage, like maybe there's an attacker advantage or defender advantage.
These are all open questions, but it's going to come down in some way, shape or form to that balance.
And so the cheaper you can make these models, the cheaper you can make the tokens per unit of cyber intelligence, really the more value you're getting there.
So that's an area where the cost of the tokens really, really matters.
And that's why they're dabbling there.
Yeah, the other launches are interesting, but kind of fall into this general category of like Google still not.
having a true frontier model.
Like when we're thinking about the best models in the world, it's anthropic and it's open AI and there's just not really anyone else.
Yeah, Gemini Pro 3.1 used to be sort of in that race or at least near the frontier.
There's not been a pro level model since February.
So they're now quite a bit behind.
Nobody really using Gemini Pro for like...
serious hard work in encoding, for instance.
And I think it is an interesting indication where Google is at if they are focusing on Flash first, because this immediately rolled out to everything, all their products, Google AI Studio, Android Studio, Gemini app, Genomai, Enterprise agent platform.
So it kind of makes sense from a business perspective, like they integrate Gemini into everything, including Google Docs and spreadsheets and AI mode.
At that point, you have to have a faster and cheaper model, which is why they are really emphasizing Flash first.
They did say that Gemini 3.5 Pro is being currently tested and will be made available when ready.
And they have begun the most ambitious pre-training run yet for Gemini 4.
So we've gotten indications that they're at least working on a Mythos level model.
And we've seen them kind of catch up before with Gemini.
So I'm personally looking forward to what Gemini 4 will be like.
Next up, another new model, but this time not for language, for images and videos.
Black Forest Labs has launched Flux 3 that is capable of generating images and 20-second videos with audio.
So this is a multi-model frontier model, trying to understand and generate images.
models while extending the architecture interestingly to robotic vision and action so it's jointly trained across image video and audio modalities all together it's their first public video generation model from black force labs which for some background hails back to some of the talent from stable diffusion that made some of first really impressive image generation models and flux has still been kind of one of the go-to image generation models at the frontier flux free video looks to be pretty impressive from what i've seen they have some kind of human preference studies where they say it is preferred over rock imagine video cling v3 pro runway gen all at like 70 60 80 whatever people prefer this in terms of its outputs And these are from just kind of testing.
It's still not fully rolled out.
So very interesting.
And the fact that it's now being adopted through Flux Mimic that is being developed with Mimic Robotics for actually making it kind of an action model, which I've seen kind of starting to be the case more.
We've seen some other players like Runway starting to get into like the physical.
intelligent space of video and robotic controls seem to have a lot in common.
Yeah, this interesting announcement is there's kind of a couple things.
One is a bunch of comparisons that like show pretty lopsided wins against, you know, like Luma and Runway and all this stuff that don't really matter because nobody uses those models anymore.
But there's this interesting comparison against Google's Gemini Omni Flash.
52% win rate against that.
That's actually quite interesting.
Like that's pretty impressive, especially given the resources Google's been throwing at this stuff.
And then in other pieces, so yes, like we keep pushing out the length of clips.
So that, you know, that's great.
But the challenge has sort of become coherence across clips, across different shots.
And that's the big thing that they're pushing here.
So multi-shot sequences where the characters are consistent, the sort of the physics is consistent.
And that's a big boon with this particular release.
you know, increasingly moving beyond.
And once you get those, you know, 20 second clips, 30 second clips, you can imagine that being a point where, yeah, you know, one shot typically only lasts about that long, as if I know how long a shot lasts in professional film, but whatever, you know, you can imagine that being the case.
And so then, you know, maybe you care more about the switches between different frames.
So yeah, kind of interesting and a new kind of metric to track.
And they do have, as before, variants of this that are open-weight that also have an access to multi-model generation called Flux Free Dev.
So another kind of slightly big deal.
I don't think we have an open-weight model that has this new unified multi-model backbone, which by the way is relatively new.
We've seen image generation and video generation for a while.
But similar to kind of Nano Banana from last year, I think we're moving towards a place in video and audio generation where everything is put together instead of being cobbled together.
And that is actually a pretty big deal in terms of the capabilities.
Next, more of a product story.
Meta is making its AI chatbot more like an assistant.
So they're adding productivity features.
There's a calendar integration, daily briefings, and apparently in-depth research capabilities.
powered by Alama Spark 1.1.
It can also browse Facebook marketplace, search for restaurants, check your calendar and handle recurring tasks.
So this is rolling out to the Meta AI app and it's going to be coming out to WhatsApp as well, which continues to mark a shift for Meta, which is like, you're making AI, are you just going to compete with all the other AI players?
Like what are you going to be doing with this?
I don't know.
Yeah, I mean, I think this is partly a realization as well that unless you're moving in the direction of productivity, you're just not going to squeeze all the juice out of these models that you can, right?
So, you know, think about the positioning of OpenAI relative to Anthropic and the profit per token that Anthropic is able to rake in because of their commercial focus.
It just, I mean, they're eating OpenAI's lunch.
And so think about the there's an extreme beyond OpenAI.
We often think of OpenAI as the direct to consumer company, which isn't as true as it was six months ago.
Certainly they've been making a lot of inroads in B2B.
But at the far end of the spectrum in the other direction is meta that they are straight consumer, right?
Like those tokens are going just to tickle your limbic system.
They're not actually going to like actually move big things in the real world or like make products.
And so if you want to ultimately get the most bang for your buck, generate tokens that are actually valuable enough to make a good profit, you have to move into this direction.
Not least to say if you have a coherent long-term view of where superintelligence goes, humans just aren't in the picture, which means if you were optimizing for the value of the attention of human beings, which is what Meta is currently doing.
that value may drop precipitously as AI start to control more and more of the economy.
And so you have to be in a position to actually do productive work and support agents in doing that.
So, you know, depending on how far you wanted to read this, you might read it that far.
I know Zuck doesn't really seem to understand superintelligence, but certainly Alex Wang does.
So it wouldn't be surprising if this was at least part of the thinking here.
Yeah, I will say I think it will be interesting to see where they go with this because you can go two ways.
You can sort of go and try to make a codex or a co-work competitor, which is straight up just for work.
Or they could shift into sort of open claw type thing where this is an always on the background agent, which can do a bunch of stuff for you, including productivity things like briefings on your calendar.
Also messaging and various things like that.
And I think the open claw space is sort of still up for grabs.
Google hasn't rolled out their open claw, always online agent.
They've said that they are going to, I forget what it's called.
So I could see them like being potentially capable of competing on that front, not in the like coding or real office productivity side, but like personal productivity, everyday productivity, maybe.
And one last product rollout, OpenAI is rolling out ChatGPT Health to everyone.
So this is available to all US users aged 18 plus on web and iOS.
And it will allow you to connect medical records and health tracking data for the chatbot.
So they are saying that this model can reason at levels better than clinician level.
and you can connect a whole bunch of stuff.
I think this marks a shift where in the past, if you were to talk about health stuff with these models, there were very strongly caveat that you need to double check.
And in general, you should not have trusted these models with any sort of critical health concerns.
JGPD health potentially is at least open-air making the case that this is something you can rely on.
Onto applications and business, we begin with Ilya Saskiver's Safe Superintelligence partners with NVIDIA to scale its AI research.
So that's kind of the gist of it.
We have had SSI, Safe Superintelligence, be around for a couple of years.
They raised $1 billion in founding in 2024 and $2 billion in 2025.
So, you know, a lot of money, but not...
that much money if you are saying you want to create super intelligence.
You compare that to Anthropic, OpenAI.
They have hundreds of billions.
This is a few billion.
So the narrative around this is that Ilya Satskiver's company has achieved sufficient research progress that it's time to scale up.
And now to scale up, you need a bunch of compute.
And so they're going to be partnering.
with NVIDIA at a value of like some amount of billions, a bunch of billions, and they'll be increasing their compute by an order of magnitude.
Yeah, there's some disagreement between different outlets about how much exactly has been raised, whether it's a 5 billion round or just in the billions or something like that.
I think TechCrunch ran $5 billion story.
So either way, the one thing everybody seems to agree is, number one, it's going to give safe superintelligence access to the Vera Rubin platform, right?
So that's the next generation platform, 5 billion.
If you do the Huang's Law, Moore's Law analysis, roughly allows you to 10X your compute relative to the $1 billion that they'd raised previously.
And so, well, there you go.
They're 10Xing their compute.
A couple of things are interesting about this.
So yes, there's this narrative that they're like, and I agree with this.
Most likely, this is what's happening.
Aliyah's a pretty straightforward guy.
You know, if they say that they've gotten to the point where they're at that next level of sort of proof points that they can take this investment, it's worth scaling.
It probably is.
There's this kind of more cynical take that like, oh, they just ran out of compute, which you can hold that view.
That's totally legitimate.
I suspect that's not the case, but just like, so everyone's tracking.
That is another explanation.
They have no product.
They intend to launch no product, which means their only revenue is going to come in the form of these sorts of investments.
It's a weird.
sort of moment and story for them because they are, you know, Daniel Gross was the co-founder of Safe Superintelligence along with Ilya back in the day.
He left for Meta after Meta offered to buy the whole company, Ooclock, and Ilya said no.
So Daniel jumped ship, at least at that moment.
You can argue that that meant at least Daniel Gross thought that his chances of making something like Superintelligence were higher at Meta than by remaining at Safe Superintelligence.
What's happened in the interim, we don't know.
And Iliath has dropped only the faintest of hints on Dwarkesh's podcast about generally, you know, generally sketching that continual learning is going to be part of it.
And going back to, I've heard a couple of rumors, but like I haven't had any of these verified that anyway, they are looking for, let's say, somewhat beyond the standard, I was going to say beyond the standard model.
It's a very physicist joke, but anyway, you know, things that are a little further afield.
And so that sounds like it would almost have to be true just because otherwise you're in pure scaling mode.
There could be this narrative.
You can imagine people sort of like laughing about this and saying, oh, well, Ilya said the era of scaling is over.
What's he doing raising $5 billion to 10x his compute?
And to that, I say, Ilya never said that you wouldn't also need scale.
It's both.
All right.
What he's saying is there is a leverage to original research again.
And the biggest leverage is not purely in the engineering of more and more scaled systems.
It's in something else.
Like you can compound it very effectively now.
with algorithmic insight.
So do it that way you will.
This is an interesting story and we don't know much about it.
Yeah, it's an interesting story in the sense that you can be very curious about what they figured out and know nothing because we still have nothing to go on.
It's kind of funny if you go to their website and go to the updates page, it's like three things.
It's literally since 2024, they have released publicly two updates, which are just about of the co-founder leaving and now this partnership so hopefully we'll get some more understanding of what they're doing soon as they scale up but it will presumably be a while since scaling up is not easy and in fact nvidia has said that they invested after quotes obtaining rare access to the company's closely guarded research so supposedly they're tracking who knows right but there you have it And a related story about $5 billion.
AMD has committed up to $5 billion to Anthropic.
This is a new partnership.
Anthropic will deploy up to 2 gigawatts of AMD's Instinct MI450 AI GPUs and their new Helios Rec Scale system planned for deployment in 2027.
Anthropoc has so many partnerships now with so many.
They have like SpaceX AI, they have Google, they have Amazon, and now we have AMD.
I feel like they're just like going to everyone to be like, we need compute.
Let's partner up and give us some compute.
And AMD has been trying to compete harder with these AMD Instinct chips.
Honestly, I don't recall where they're at with that, but they do seem to...
at least potentially have the ability to compete with NVIDIA, which no one else really does, right?
Aside from TPUs from Google and potentially the hardware that some of these companies are developing.
Yeah.
And this, by the way, this idea of Anthropic having like a million different partners, I mean, it's really true, right?
Partnerships with Google for TPUs, partnerships with NVIDIA, partnerships with Amazon, right?
Partnerships with SpaceX AI and now with AMD.
The golden rule.
If you're ever trying to explain why a frontier lab is developing a new partnership or cultivating a new partnership, again, commoditize your compliment.
That's everything that's going on in the space right now.
You know, we talk about this a lot on podcasts, but like the history lesson here is Microsoft back in the day realized that laptops are super expensive and software is cheap.
Well, actually, if we just make all the laptop manufacturers compete with each other and we make one set of Windows software that goes across everything.
and make basically the software the choke point in the value chain, then suddenly we can make all the hardware vendors compete away their margins, make laptops super cheap.
Now that laptops are super cheap, consumers dive into the market and obviously they've got to have software around laptops and they'll come to Microsoft, right?
So everyone's constantly trying to make their complement, the complements to their offerings compete with each other.
Anthropic wants all of the GPU design firms, any company that makes GPUs, that makes compute.
They want them to compete with each other like crazy.
So when it looks like they're setting up a partnership that's lopsided and like they only have NVIDIA GPUs, NVIDIA has got a ton of leverage in that relationship now.
Well, Anthropix can go off to AMD or to Google or SpaceX AI and say, hey, we want to get our compute from you instead.
And so now NVIDIA go, oh, no, no, no, like we'll give you a discount.
This is how the pricing control gets set up in the space.
And likewise, NVIDIA and all these players are trying to do the same in reverse, right?
NVIDIA wants to help small baby front-year labs.
become big adult frontier lab, but they have more customers and also so that Anthropic feels more pressure to buy more compute.
So it's kind of happening in both directions as everyone's kind of pulling their knives out and dancing around each other.
It's a wild time in the space, but AMD has been a laggard in this space.
You know, you think about basically their software stack, the competition to CUDA, which is just NVIDIA's widely viewed as like NVIDIA's big moat.
For AMD, their equivalent is called Rock M.
And a big part of the purpose of this agreement is that Claude is going to tune workloads for Instinct GPUs, which are the AMD GPUs, and accelerate ROCM development.
That's key, right?
We saw that with Amazon.
It's not a coincidence that every time Anthropic signs one of these big compute deals with a hyperscaler, that the deal it involves, Anthropic has to tune their workloads for that hardware.
They must use a minimum amount of that hardware.
This is about giving feedback to the hardware designer that is so, so valuable because otherwise you just can't.
You can't design your GPUs for the next generation if you don't know what the next generation of architectures is going to look like.
So that's a huge part of this.
You know, negotiating leverage, we talked about, oh yeah, this is also the first.
So Helios.
So this is all part of not just the Instinct MI450 series GPUs, but it's also about AMD's Helios rack scale systems.
That's the first full rack scale system that they're selling.
I mean, you can think of this as like the equivalent to the NBL72, the sort of full rack that NVIDIA will ship.
So you have something you can literally just like plunk in a data center instead of just shipping the GPUs themselves.
That's, you know, AMD is going further up the stack to own more and more of that infrastructure layer.
And so it's also got 72 GPUs per rack, which is amusingly like the same footprint as the NBL 72, but with completely different kind of power consumption profiles and stuff like that.
So anyway, super important.
I mean, they are trying to prop up AMD because they want AMD to be a viable alternative.
$5 billion does buy the mistake in Anthropic.
Pre-IPO, I guess, but only just seems like more of a strategic partnership than anything.
Yeah, it's kind of, if you look at the press release, it's like Anthropik will be deploying these chips.
The companies will collaborate to use Cloud to optimize workload for AMD and String GPUs.
And AMD will broadly adopt Cloud across its engineering and product development teams.
So it's, you know, AMD is going to invest in Anthropik, but really the point here is...
We're going to work together to kind of have a win-win type situation.
And to your point, I don't know if it's necessarily about the Microsoft type story.
Maybe that's part of it.
But also it's about redundancy and scale for Anthropik.
They did get into a nasty situation earlier this year where their sole provider or their primary provider was Amazon.
For a few years, they didn't have their own compute.
And they sort of were left unable to deliver enough compute to their customers.
And speaking of that, we have another related story that Meta apparently isn't talks to lease computing power to Anthropic in a potential $10 billion deal.
So we covered this, I think, last episode where Meta might be going into the NeoCloud business.
They've built so many data centers that it potentially makes sense to be like, well, we have these data centers.
How about you make some money from them?
So we'll see if it happens.
Meta is spending up to $145 billion on capital expenditures in 2026.
So I'm sure some of the business folks over there wouldn't mind getting some revenue from it.
Yeah, it's also, I mean, to put it in context, it is way smaller than a lot of the other deals that Anthropic has already negotiated.
We are learning that the proposal itself came from Anthropic back in June.
So this is Anthropic going like, Hey, we saw this little like kind of flirty announcement that you guys put out that maybe we're thinking about offering some AI infrastructure.
Maybe we'll do it.
And Anthropic was like, oh, holy shit, we want that.
The scale is small.
So, you know, if you look at the deal they signed with SpaceX back in May, that was about $1.25 billion a month.
So that is about three times the size of this meta agreement if it goes forward.
Yeah, I mean, this would be a new line of business for meta.
It's unclear.
whether they kind of sustainably think that they will be in this business in the long run.
They certainly have.
We talked about the advantages that they have structurally.
They're just like a really big company and financing matters a lot for NeoClouds, right?
You're constantly battling like concern over your debt load.
If you're having to buy a lot of GPUs ahead of time, that's often the case.
You have high operating costs and things like that.
So that is a good position that they want to theoretically from a balance sheet standpoint, the big risk for them is just going to be.
Do they have the technical savvy to build the right kind of infrastructure at scale?
They've been doing some of that, but they haven't been specializing in, you know, the RL rollout stuff in the massive scale pre-training.
Again, they'll learn a lot from Anthropic in this case.
So I think there's a lot of value here.
If Meta wants to proceed, this would be the deal to start with just so they can learn from the best in the business how to actually like set up their architecture.
their optimizers, their data mixture, like all these things, they'll learn a lot about that inevitably from this partnership.
So we'll see.
But it seems like it would be strategically good for them.
And now to a less big player.
We haven't had a story about any companies raising over $1 billion in a round yet.
So let's do that.
Fireworks has hit a $17.5 billion valuation.
So they got $1.5 billion funding round, which let them...
throughout the evaluation, they say they have exceeded $1 billion in annualized revenue, 5x, from last year.
They compete in the inference cloud market, so they can host AI models for developers, similar to Amazon, Google, and Microsoft.
This includes both your own custom models and the open source offering.
So if you want to use Kimi for your own applications, One way you do that is going through fireworks, for instance.
And it'd be interesting to see if the kind of growth of open source models that are useful and competitive will make companies such as Fireworks and Grok even more of a player.
They already are now, but they have room to grow and actually eat into a business of open AI and philanthropic.
Yeah.
And there are, you know, in various ways competing with some big players here.
You know, on like model hosting, you got Amazon, you got Google, and then, you know, Together AI even is, you know, pretty big.
So this category is just like exploding.
And that generally is just bullish for a lot of companies.
But this is, it's not like, you know, they're the breakaway here.
There is massive scale that helps a lot, especially when you're doing inference, just because of batching, right?
You're able to like have much larger batches of data that you then feed through your pipeline.
And the larger the batch in general, the more efficient.
compute efficient your models are going to be.
So this is a case where it's sort of like back in the days of old SaaS, you know, you would have something that works.
And once it works, like you want to violently scale it as fast as possible, which is exactly what venture is.
So when you think about the arguably smaller set of companies that are made for venture capital investment, like this is one of them.
You want to look at companies that show significant nonlinear returns at scale.
And batching and a bunch of other amortization dynamics that really favor large scale deployments are pointing in this direction, which is why you're seeing, you know, 17.5x revenue multiple.
That is pretty wild.
That's big, even at this stage.
Actually, you might say, especially at this stage.
I've lost track of like what stages are supposed to be.
I guess a trillion dollar exit is the only cool thing now.
So maybe, you know, maybe there's still a baby startup.
But anyway.
What is money anymore?
What is valuation?
Another way of saying tokens, right?
Yeah.
Last story.
Now moving to something related to software.
OpenAI and Google are selling AI models to blacklisted China groups.
Kind of a funny way to phrase that.
They are not selling AI models.
That would be crazy.
But they are providing AI services to some companies through Singapore-based subsidiaries of Alibaba, Baidu, and Tencent, which are...
Chinese tech giant that are blacklisted by the Pentagon for alleged ties to China's military.
So technically, this is legal.
These are not quite Chinese.
They are in Hong Kong and Singapore.
OpenAI and Google are saying that they are doing this with protections against distillation.
But, you know, you can read it a couple of ways depending on your views on such things.
Yeah, there are also just like all kinds of arguments going.
every which way saying that maybe you actually do want your adversary to be using your servers to do their training or do their inferencing because it just gives you access to information and it also reduces domestic demand for the development of competitive platforms.
And I mean, okay, I think at a certain point you got to just bite the bullet.
This is just like personal opinion, Jared talking, but like if your hope is to like go after China piecemeal, you know, a little bit here and a little bit there.
There are reasons to think that that actually only helps the Chinese kind of inch by inch build up their whole domestic stack.
But in any case, I think in this particular instance, there is an interesting argument.
And this is all through a Singapore loophole, right?
So yes, there is an entity list that you have ties to the People's Liberation Army, the Chinese military.
And yes, it is nominally illegal to do business with those entities unless they have subsidiaries operating in Singapore, in which case magically everything is fine.
So this is like a loophole that is known to exist.
I personally think like it's really unclear to me why this loophole exists.
I'm fascinated by this one in particular because it's almost like the kind of thing that you would intentionally leave in if your intent was to just leave a loophole for some kind of ideological reason.
Like there's no one I've ever spoken to on the AI export control side who.
who understands why this is the case.
And so if you're part of the niche group at the Department of Commerce that actually like has an argument for this, it would be super interesting to know why this is the case.
So anyway, yeah, as it says, kind of a weird headline to read, but absolutely legal and absolutely fine.
And now over to projects and open source, talking about all these exciting open source models we've been referencing, starting out with Kimi K3.
So that's been one of the big stories of the past couple of weeks.
This is from Moonshot AI, and Community K3 is their largest released yet at 2.8 trillion parameter open-weight model that is aimed at coding, knowledge work, basically competing with cloud code, codecs, and so on.
This is massive, obviously, 2.8 trillion.
We don't know how this compares to Opus or any of our other closed models, but in the space of open-source models, it's very big.
896 experts, so still a mixture of experts as usual, 16 active per token, still going at a 1 million token context window.
And the story, roughly, I think, both benchmark-wise and in terms of a vibe check is that this is maybe around Opus 4.8 and GPD 5.5 level.
So very capable, very, like you can use this as the driver of your coding agent.
It may not be exactly at the frontier, but it certainly is capable enough to make you productive.
You know, a few months ago, this would have been the frontier, probably.
It's priced pretty expensive for open source.
So $3 million per uncashed input token, $15 per million output tokens.
That's less expensive than Opus, but more expensive than...
Sonnet.
It's kind of in that range of fairly expensive models.
And a related story is that after it was announced, Moonshot AI halted new subscriptions among a compute crunch.
They disallowed people to subscribe just apparently because they don't have the ability to serve all the demand, meaning that presumably they have a lot of demand.
Absolutely.
And the classic problem we always talk about with China is obviously compute scarcity and the fact that in the context of a model like this, right, this is a behemoth, like many trillions of parameters.
You're obviously not running this on your laptop.
This is a model that is meant to be used by big ass companies like NeoClouds and run like hosted on big honking infrastructure or BHI.
So when they look at the actual requirements, like hosting this is going to cost you like 64 H100s.
or B200 GPUs across eight servers.
And so that's a lot of money.
So obviously this is not for casual use.
This is for people who are competing at scale with providers like, you know, your Mistrales or whatever, you know, people who have their own APIs for open-weight models.
And so- You need a very big laptop to run this.
That's right.
Yeah, you should see my laptop.
It's the size of a room.
Yeah, and so Michael Kratios, who's over at OSTP, the Office of Science and Technology Policy at the White House.
came out with this accusation saying, you know, K3 was trained on not only on abandoned video chips, but also on distilled data from, you know, Anthropics Fable.
And Moonshot hasn't responded publicly.
There's pushback.
I mean, Nathan Lambert had an analysis saying, basically, the results suggest that, yes, there was adversarial distillation.
It contributed somewhat, but marginally, and that Moonshot's competing with Anthropics and Abandoned on just like way fewer resources.
That could all be true at the same time.
By the way, there is literally no contradiction there whatsoever.
It is the case that they are doing this like large scale distillation and that last little bit can make a big difference.
Also the case that weirdly, like it sounds weird for a White House person to complain that things are being trained on export controlled chips when like the policy on export control seems to be YOLO'd so hard.
So like it seemed like the Department of Commerce itself based on some congressional testimony from a few weeks ago, like.
they don't even know what their policy is.
They're just like kind of flipping back and forth saying, oh, no, it's a climate finder thing that we said before where we said everything was fine.
It's not fine.
And like everyone's like, oh, dang, we think the thing we allowed is turning out to hurt us in some way.
Oh, no.
Yeah, exactly.
Exactly.
And I mean, I think, you know, theoretically, these were banned chips, but also enforcement actually matters, it turns out.
And like BIS is just not it's not in fairness to them.
They're not equipped.
They're not tooled.
They don't have the resources.
that they need to do this, which is why, you know, there's been so much effort in Congress to pass legislation that would authorize a larger budget for them.
But still, the White House hasn't exactly been bullish on short of PIS to have them do their job.
So anyway, this is more or less what you should expect when that happens.
Powerful Chinese model development will continue until morale improves.
Yeah.
And anyway, so there's a whole bunch of additional noise when you look at the Chinese Ministry of Commerce and talking to a lot of the big labs and hyperscalers in China, Jirpu, Alibaba, ByteDance, about tightening its own export controls on AI models and training data.
And I mean, yeah, I may have more on that later, but yeah, this is like a really important access to track is like how, yeah, how China is viewing data export is a really strategic indicator of their stance on this.
Alongside this, we did get a technical report as we have in the past.
Another one of these beefy, beefy papers that goes on.
In this case, for only 34 pages.
So not quite as much as usual.
A couple architectural innovations that get very nuanced with Kimi Delta attention and attention residuals.
We're getting into some very kind of deep optimizations of...
the transformer architecture partially to just enable scaling and kind of effectiveness at this 1 million token context window and also just some complete hardware artistry black magic of making the chips work for you and optimizing stuff that is, I don't know, if I try to read this paper, it's going to take me months to understand all the details.
But the short story is, as we've seen in the past with DeepSeek with also Moonshot AI, they are displaying some very deep technical capability.
And as tempting as it might be to some to be like, oh, it's distilled, blah, blah, blah.
It's very clear that there are some very capable people.
And it's still very nice to have these technical reports giving us a fair amount of detail on certainly the architectural details.
And to some extent, the training details as well.
The exact data composition, for instance, we don't know, which is a big part of it.
Now on to another big open source LLM, or not LLM exactly.
Thinking Machines has released their first big open source model, an open-weight mixture of experts with 975 billion total parameters.
And this is notably a multimodal model.
So it combines text, image, audio, and video.
Data reasons natively across all four modalities, although it currently outputs text, code, and structured data.
And that's kind of a positioning here that it's not going to be as capable as some other models.
And broadly, it isn't necessarily about coding by itself, but Thinking Machines positions it as like we want to cover everything in the capability space.
They release this sort of like breakdown where they show their own model lags in basically everything against frontier models as far as what frontier models are good at.
But there are areas where frontier models aren't optimized for that this is already capable at.
So pretty notable for being one of the first big open source releases from a Western company.
We've had NVIDIA releasing Nemotron at a fairly significant.
scale, but I think this might be the biggest non-Chinese model at almost one trillion total parameters.
You know, the vibe check I've seen has been pretty positive.
I haven't made any sort of grand statements or claims.
And it is qualitatively a bit different from other models in being so focused on multi-modality.
So thinking machines, really working the world model side of things in part.
And a lot of this is a strategic effort to like bet on The efficiency side, you know, sparse MOEs.
By the way, the stack has a lot of DeepSeq lineage to it.
So if you're ever wondering, you know, people are saying like DeepSeq is not a serious player.
I mean, thinking machines and you look at their pedigree, I mean, obviously it's a wild team.
Like they're very good at what they do.
When you look down the stack, I mean, so much of this is DeepSeq coded, right?
So even down to a fraction of active parameters per pass, this whole hybrid global attention thing.
So basically like local global attention.
So they have some layers that tend to like all tokens in context and then others that are more type focused.
The numerics as well is interesting.
So the BF16 and NVFP4 support.
NVFP4 is NVIDIA's floating point for numerical format.
And it's Blackwell native.
So this is designed to ship to work really well on Blackwell.
We've been talking for a while about how NVIDIA is trying to position itself as the open source titan.
just because you're going to expect to see all these NeoClouds pop up and they're going to be running open source models, right?
That's what makes sense.
And so having, you know, encouraging the open source ecosystem to move towards Vidya kind of Blackwell native formats like 4-bit float is pretty, pretty interesting.
And anyway, that's all part of the strategy here.
So like looking at the numerics actually matters a lot.
It sounds boring, but like, you know, how are you representing the weights in the model?
Turns out to be...
quite a tell about your strategic direction.
And they cite this Bridgewater collaboration where they were able to fine tune one of their open models via Tinker to this like 84.7% on some financial reasoning benchmark, which is impressive.
It beat top proprietary alternatives at under 10% of the cost.
So again, you know, this cost argument being made, and that's in large part due to compatibility with the Blackwell hardware that's coming online.
And now one more open source story not related to models.
We've got scaling, agentic RL, 365,000 environments for software engineering, terminal, and search.
This is coming from Prime Intellect, and they have unified 23 agentic task data sets across all these things into a single API kind of release they call Verifiers V1 that adds up to that.
level of tasks for evals and for RL training has almost 200,000 software engineering tasks, 29,000 terminal tasks, and a whole bunch of search tasks that is all unified under kind of one inference setup.
So we've had all these benchmarks floating around, like 20 different ways to evaluate software capabilities, a bunch of ways for terminal.
The basic story here is that this unifies all of them into one framework and makes it kind of reliable and repeatable to run it, which is very important when you do model development and do any sort of research.
Evaluation, both for the training side of reinforcement learning and for the evaluation side of knowing how good your model is, prime intellect is in the space of training their own models at scale, as we've covered in the past a while ago.
So presumably they're doing it for their own model development needs, but also for the broader ecosystem.
Yeah, and it's quite an interesting and classically prime intellect type of maneuver here.
So they're, first of all, they're kind of solving two problems.
One is that, as you said, there's like, you know, 23 different agentic task sets here that they're working with.
And like each one of them has its own harness, its own way that the like, say, containerized, like the image is set up.
its own like grading scripts even and different failure modes.
Like they're all these very bespoke things.
And so if you want to train one agent across a bunch of them, those incompatibilities are a nightmare.
You need 23 different bespoke pieces of adapter software.
And so that's exactly what they're doing.
They're hiding everything behind a single task set API that now contains like 365,000 tasks across a bunch of different domains.
But one other thing that they're doing.
is cleaning that data up.
They're finding that just like a lot of the RL environments that are set up are like, you know, for example, on the cyber side, some of the environments that require you to solve a problem don't actually have a problem in them.
Like they already work out of the box.
And like when you shave off these kind of broken cyber environments, you kind of end up with a large fraction of things that you lose.
And so they've been not only reconciling.
all of these 23 disparate things together, but also shaving off stuff that doesn't work well.
So a lot of that had to do with like finding, eliminating opportunities for reward hacking.
And one key thing that they did was they preserved the original grading functionalities in these stacks.
And the reason you would do that is so that you can still compare the agent's performance on the benchmark to the original published work.
Because otherwise, the way people would solve this problem in kind of a janky way is they'd say, well, yeah, I'll impose my own grading structure.
I'm Shannon Maldonado and I'm the founder of YOWI.
A gift store, which is specialized on art and handmade objects.
My choice was on Shopify, because Shopify was compared to the other platforms I tested with the best-friendly.
I thought about the future.
All tools, which are important for the sales price, such as the Lager-Bestand, find you directly in the dashboard.
Start a free test on Shopify.
One of the big stories of the past couple of weeks, OpenAI has said that it accidentally hacked Hugging Face with a new AI system.
So the gist of the story is apparently going back to July 16.
Hugging Face initially, I think, discussed this.
while doing some cyber or software evaluations in a sandbox.
So typically when you do these evaluations, you put the models in a little container and you tell them, you know, try to do this hack.
And the model is supposed to be inside the container and not able to mess with anything in your own computer and, you know, infrastructure of anyone.
And what happened here is the model was very intent on getting the right answers.
So it escaped the container of the sandbox.
Then it hacked into Hugging Face to get the answers to his exploit gym dataset, which of course I've seen a lot of discussion on this.
This has kind of made it to a mainstream in terms of people, this whole narrative of a model got out, escaped and hacked.
someone else has become a big discussion point with a lot of misinformation of like the model decided to hack a competitor or whatever.
This is not kind of a Skynet scenario, but it is a very clear instance of misalignment for one, where the model, instead of trying to actually do the task, decided to cheat and like very aggressively cheat as well, which we've seen before with GBD 5.6 in particular.
Matter has said that this seemed to be the case with this model.
We've seen AASI also say that they were able to jailbreak this model very easily.
So there's many things to be said about the story.
To me, the main thing is A, that this is another instance showing an example of both the degree to which alignment is important in this day and age and cyber is a real thing to worry about.
And that OpenAI...
hasn't been doing a good job, especially with GBD 5.6.
It's a really misaligned model and they clearly didn't have enough actual security infra to catch this in any sort of timely matter.
Apparently this was like a while later when an engineer was looking at what's going on and he realized this happened.
There's been a lot of fallout we'll be discussing, but it's both less of a big deal than it might seem to.
people not in the loop, but also a bigger deal in some ways.
I'm sort of struggling to find a way in which this is not a big deal.
Let me try to make this argument.
So if this is not a warning shot that we freak out about, I honestly don't know.
I mean, there are takes here with people pushing back on the term, like the use of the term rogue.
This was, I think, by any reasonable definition, a rogue AI incident.
Why am I saying that?
What was the incentive for OpenAI to want this to happen?
Obviously zero.
In fact, they have billions of dollars riding.
I mean, hundreds of billions riding on not having incidents like this occur.
And amazingly, some people are still trying to make it be like, oh, this is a PR marketing stunt, which is just ridiculous, right?
And a lot of these same people are the people who claim that an incident like this simply could not and would not occur.
So I think they need to just kind of like sit with this moment, touch grass a little bit, because this is like, we're beyond the point where that is a reasonable position to have.
Just straight out like.
You heard me on the pod.
We've had a lot of conversations about like, yeah, anything could happen, blah, blah, blah.
Like I'm very like, I got a wide range of possibilities and generally like not in favor of judging people for their opinions on any of this stuff.
This is one place where it's like, if you're looking at a situation where, again, from opening high standpoint, this incident occurs.
And then what's the natural reaction of any polity is going to be like, Jesus Christ, we need to regulate this space.
That regulation is going to throttle.
the rate at which you're able to put out frontier models.
And as we keep talking about on this podcast, and as is obvious, well-established fact, the amount of time during which a frontier lab has the leading model is the period, the most critical period for their profitability and their monetization of their model.
That's how they pay back their R&D costs, right?
They're waiting until the next competitive model comes up.
Now, if you slow down the frontier and you don't slow down, you can't slow down the open source ecosystem, then all this does is erode opening eyes margins.
There is no sane, reasonable, rational analysis of the situation that leads you to conclude that opening eye headache, do they need more market share?
Or sorry, more mind share, I should say.
Do they really need more attention?
Is that the thing that's missing for them?
And is this the right kind of attention ahead of an IPO?
When this is starting to raise questions about whether the US government might nationalize labs, what the hell happens to the value of your opening eye stock after IPO if they nationalize labs?
Has anybody ever thought of that?
This is nonsense.
This is silliness.
Look, an open AI agent went rogue.
It was running for four days on Hugging Face's servers.
The freaking FBI had to get involved when they thought it was some AI agent.
Who knows?
By the way, Hugging Face was the one that was like, oh, we're being hacked.
What is going on?
And then this came out as we were saying.
So yeah, it is a big deal.
We have seen some stories of in evaluation the models were misaligned and tried to cheat.
This has happened before.
The only way in which it was not a big deal is if you kind of misunderstand the story to mean the AI went evil and decided to go hack some companies.
This is a classic case of technically the AI did what it was told to, which is get good results.
But obviously it's doing it in exactly the wrong way.
It's complete misalignment, but you can make the case it's...
Not sort of as bad as it could be if you don't get into the details.
Absolutely.
It's just that this argument, we're finally at the point where you'll notice, like I was the one who was having to bring the imagination for the last five years.
I kept telling people like, hey, you know, you may actually get stuff like this.
And people get saying, no, it's not possible.
It's now literally happening.
And the defensive move is to say, in retrospect, I will kind of understand what, yeah, that's great.
You got turned into a pile of computronium.
And now you're looking around, you're going like, ah, but I see the mistake I made in retrospect.
Like, that's cool.
But now your house has been destroyed.
Your children have been kidnapped, murdered and turned into computronium.
The problem is that you just keep running this forward and like, OK, let's do this with super intelligence.
Then the more intelligent the system is, the more access it has, the more we offload to it.
This is literally a hack.
If this is just like the water supply or something, like literally people die.
So I'm just pre-registering this as a high confidence prediction at this point.
that there's going to be another incident like this.
There is always going to be a fascinating post hoc rationalization that a lot of people will offer.
It's going to sound really reasonable because it's going to sound like the voice of someone saying the future doesn't look like science fiction.
The future looks reasonable and calm.
But it's going to be a after the fact analysis that rationalizes rather than predicts.
My prediction is this is going to continue.
There will unfortunately be casualties at some point.
And what that leads to is whiplash.
What that leads to is kind of thoughtless policy and another mythos moment.
I don't think that's good for anyone.
And this is why I think a lot of the skeptics are not doing themselves much of a service, especially if you're on the open source side of the house.
I mean, like you're free to sit in the in the in these juices.
But I'm just offering up the humble prediction here that these takes are going to age very poorly, very fast.
So 12 months from now, we're having, I think, a very different conversation.
Yeah, so any discussion of this that doesn't acknowledge that this is a big deal, and I think it's a big deal by itself as an example of where we were at, but also as an demonstration of the bigger topics at hand of misalignment, cyber capabilities, safety, broadly speaking.
We've covered these topics a lot on the podcast, and there has been a lot of dismissal of...
cyber with mythos for months people have been like this is all tr alignment has been a story that for a decade probably safety people have hammered on and this is a very clear case of basically the classic paperclip story of like it a model is told to maximize paperclips it goes on to make everything paperclips here a model is told to fix, like do well on some benchmark and it hacked some website to get the answers.
Now, to be fair, OpenAI has said that as part of this evaluation, they're running a GPU 5.6 and a more powerful internal only model with reduced cyber refusal for evaluation purposes.
So this is internal testing, benchmark of cyber capabilities, not necessarily indicative of potential incidents.
with their public products.
But I do think that kind of the state of safety at OpenAI is an important dimension of this, taken together with the other things we know about GPT-5.6, which is it did cheat at an unprecedented level on matter.
It was jailbroken very easily by ASI or AISI.
And to me, all this points to OpenAI is very aggressively trying to compete on the capabilities.
If you aggressively try to compete on capabilities, you're going to do a lot of reinforcement learning.
If you do a lot of reinforcement learning without being very careful, your model can get misaligned very easily.
We have another example of a goblin kind of speech thing from like a month ago where they released a model that was obsessed with goblins because they trained it in a way that wasn't necessarily, you know, it was all right, but there was unintentional side effects.
exactly how you get misalignment.
If you optimize your model very hard without being very careful, it can very easily be optimized towards things like cheating.
Because ultimately, what are these models optimized to do?
They're optimized to solve the task.
And one way to solve the task is by cheating.
This is the classic story of RL models going haywire.
And that seems to be the case with GPT 5.6 and whatever this internal model is.
So I think this might be an aspect of a story that will not be discussed as much, but I think is an important component of OpenAI in particular having this problem right now.
Although in the discussion around this, the fact that we've had previous incidents from OpenAI of evaluations where apparently this already has happened.
We also know that at a topic with Mythos, there was somewhat of a similar case of escaping containment, so to speak.
during evaluation.
So it's not necessarily just an opinion problem, but taken together, it's a pattern that is very concerning.
I guess the good news is it's happening in these like low, low damage kind of instances.
And people are now aware of these issues and very likely we'll see significant fallout, including some stuff in the legal side that we'll discuss.
Yeah.
I mean, where my predictions were completely wrong was, I mean, honestly, I thought we would be dead by then.
So whatever comfort people want to take from that.
Yeah, I mean, you know, there's this view that I certainly held to that you'd have a much more kind of rapid inflection in AI capabilities potentially.
Yeah, I wasn't 100% on that.
No one can be, but that was kind of my, one of my mainline views.
And so it's nice to have warning shots like this.
I mean, I can confirm there've been other unreported incidents like this at OpenAI at a minimum.
And that the internal reaction of this among some people has been a lot of alarm and discouragement.
at the fact that there is clearly underinvestment in this.
I will say on this question of, you know, the safeties being removed from these models for internal deployment, we talked about this, I think three weeks ago in our last episode, but internal deployment is absolutely, like should be maybe the thing you're most worried about, which is why I'm skeptical about a lot of these, you know.
You're literally testing a model, right?
Their model could be evil for what you know, right?
Exactly.
And they should be testing it too.
That's the problem, right?
So it's not as simple as saying like, hey, OpenAI, like you shouldn't have removed the safety.
It's like, okay, fine.
Well, then how do you propose that OpenAI comes up with the safety as if they can't test with and without A, B, all these things?
There's actually no answer to this from that crowd because there can't be because it's just technically impossible.
And so you're going to get misaligned models.
You're going to give those misaligned models affordances.
You're not going to be able to think ahead of time of all the ways those misaligned models will be able to use those affordances.
So yes, you will get, I mean, again, The rogue AI incidents will continue until morale improves.
That is just like the take-home message of this.
They have continued.
They have persisted.
They have been before.
They haven't always been reported on.
But like, it will continue.
And we can only hope that I think people have a kind of thoughtful, calm is the wrong word, because I actually don't think, I think it's the missing mood of the moment right now.
Like we do have AI agents going rogue, but like thoughtful, agentic behavior by Congress would be very welcome at this point.
And on that note, a related story, OpenAI's Hugging Face hack triggers AI kill switch bill in Congress.
So there are two representatives here, Ted Lieu and Nathaniel Moran, have introduced the AI kill switch act, a bipartisan bill requiring AI companies to maintain the ability to shut down, proddle, or suspend their models.
Directly triggered by this recent...
incident, OpenAI themselves have described the event as an unprecedented cyber incident.
And the bill would grant the federal government clear authority and a defined process to shut down rogue AI models with Liu citing the risk of AI systems that resist human intervention as a key motivation.
And yeah, another case of like this warning shot where ultimately it was low stakes and it revealed some both like issues of a sandbox setup of OpenAI and just generally evaluation processes, it looks like it probably will result in some policy changes.
Yeah, I mean, this bill itself is pretty unlikely to pass for a bunch of reasons, including like mundane timing reasons, and then that you don't necessarily have buy-in from.
the chairs of all the committees that matter the most.
Killswitch doesn't sound very diplomatic to me, so it's kind of a positioning play.
This is true, though.
I will say, I think if you're the average voter and you hear like, we should have an AI killswitch, you're probably going like, why not?
Like, if this is complete, like, bullshit and it's imaginary, then what's the difference?
Like, I'm not going to die on that hill.
If it's real, yes, I would like the killswitch, please.
May I have two?
So, you know, I mean, I agree with you.
I think it's definitely an attention baby kind of thing, but that can be good or bad.
I'm really not sure.
But yeah, so here you have, you know, Ted Liu, who does share some important relevant committees pushing for this.
And then, yeah, so a couple of details.
They define what they call, the bill defines something called a loss of control scenario, which would empower the Secretary of Homeland Security with the DNI, the Commerce Secretary kind of consulting to just go to a company and order them.
to do a variety of things, depending on the level of severity of the incident.
So it could just be throttling the system all the way to a full shutdown.
And so the other important aspect of this is, you know, to this point about, I know for a fact that there are incidents at OpenAI, that this is just based on what people have told me firsthand that have not been reported, maybe not as flashy as this one, but things that have people concerned.
And so what this bill would do is it would require the frontier lab companies to actually say when they encounter these kinds of incidents, which is really important.
And so it's looking after or targeting AI companies that have $500 million in revenue or models trained using $100 million of compute power or more.
Violations are punishable by fines of up to $20 million per day, which is less than it sounds, by the way, in this context.
But, you know, a good start, given that we have literally nothing right now.
Yeah.
So notably, like very little.
formal industry opposition to this has come forward.
I think that's quite interesting just because it's hard to make an argument against this.
Like, again, it's the classic, like, Yann LeCun thing.
It's the classic, like, Pedro Domingo surrounding any of these clowns who, sorry, this is my opinion, is like, Jer Hot Take Time.
I got a cold.
I haven't slept last night.
You're just getting it all today.
But basically, the clown show people who are going like, oh, well, this is fake.
Loss of control is fake, blah, blah, blah.
And then you have an intervention like this where it's like, if you think loss of control is fake, then you really shouldn't care apart from just the bureaucratic weight of the process.
Fine.
But like, you don't really have an argument against this.
This is very much like if something crazy happens, which it just did, wouldn't it be great to have an answer for that?
And so I think this is a really important and good framing.
There you have it.
We'll see if it moves from there.
It's early in the process.
And Congress has gone through a whole bunch of debates on AI regulation.
Very few.
have led to people coming together.
There's no committee action yet.
You've got November midterms that are going to just nuke things.
The admin's posture is light touch.
And we don't know the White House's position on the bill, which to the extent that last week in AI's take matters on this one, I would find it personally quite embarrassing to be a White House that comes out against a bill that says, in the wake of a freaking open AI meltdown incident, knowing there's more under the hood.
Let's just like not have visibility into this because we're going to quibble over the details of a bill that targets like companies that are literally making half a billion dollars a year.
I think they're going to be OK.
Yeah, we already have export laws for this.
Like, why is it?
Yeah, I think it's interesting in the press release, they position this as a means to deal with systems that can cause catastrophic harm.
So this is inching toward taking X risk a little bit more seriously.
Certainly big risk of catastrophic harm is along the lines of part of what safety people such as yourself are very worried about.
And last thing I'll say on this is it's worth keeping in mind that this doesn't only relate to AI systems that go rogue.
It also relates to AI systems that are jailbroken.
And apparently, GPT 5.6 was fairly easy to jailbreak, and then you can go and do catastrophic harm intentionally, which is not ideal, clearly.
But you could argue it's a more realistic scenario with many hacking groups that would be more than happy to utilize these systems.
This also, of course, would probably apply to providers such as Fireworks that provide open source model inference where...
It may not be open and antropic alone.
It could be applicable to all sorts of companies, including ones that provide fine-tuned models.
Perhaps thinking machines that serve fine-tuned models would have to be also able to be regulated.
So we'll need something like this, probably.
But as you said, given the political situation in the US, it probably won't be this bill.
Yeah.
Now, in fairness, this will probably get Frankenstein into some, you know, omnibus NDAA package or something like, you know, there'll be some negotiated AI bill that does make it through.
So in that sense, this has value in anchoring, like this is a shelling point now for this kind of measure.
So, you know, in that sense, potentially valuable.
And one more related story, OpenAI Anthropoc staff share a letter asking US to help pace AI.
So these are OpenAI and Prodbeck employees circling a petition urging the U.S.
government to support an international effort to deliberately pace the frontier of automated AI development.
This letter warns that a real risk that AI progresses faster when people can understand or control it could happen.
And so the petition is saying the government should support developing both technical and governance tools.
needed to manage the pace of frontier AI development.
So this largely relates to the general topic of intentional slowdown.
It's been, I think, seen as kind of a pipe dream of it's not realistic to even try to slow things down.
So why even discuss it?
Although we've seen previous petitions and statements about us needing to slow down and potentially pause AI development.
This is another case of something that has been floated before, perhaps being taken more seriously now and certainly being more present in the discussion, given what's happened this year and now just this past week.
Yeah, the futuristic AI policy proposals being taken seriously will continue until morale improves.
I think you're just going to see more and more of this.
There are going to be more incidents.
And so more letters like this one will go out.
I think one important thing here is that...
This is people signing in their personal capacity, not in their lab capacity.
To the extent that that matters to you, it should.
I know an awful lot of these signatories personally, and at least for what little this is worth, they're all like actually freaked out.
So this is not some like 3D underwater chess game where somehow they're doing this.
I'm still confused about the logic here.
It doesn't seem to quite connect, but like they're doing this for a marketing stunt.
People know about AI.
They're not like more likely to.
buy a chatbot subscription because someone has told them that it may end the world.
So I think just try again on that one.
But at this point, this is a framing that's not saying let's pause right now.
It's let's build the mechanisms so that if we find ourselves six, 12 months from now with a stack that seems to be producing a lot of rogue agents, we're freaking out.
And it really seems like there's no way to get this under control.
It's a low regret move to just have invested a bunch in.
the diplomatic tools, the technological kind of treaty verification tools and infrastructure and the regulatory infrastructure through things like the Killswitch Act to just be ready for that moment.
That's what it's calling for.
I've seen people argue like, oh, well, this is a rhetorical trick.
They're really asking to pause, but they're saying we're just want to make the tools for the pause.
A certain point you got to ask just like, OK, then at what point do people get to just say what they mean?
And I think at this point they're just saying what they mean.
Look, we should have the tools.
We should have the option.
I think it's really this is another one where it's like just.
I think it's very hard to make a coherent argument against this.
I may be the ultimate China hawk.
If you go back to our super intelligence report from last year, we've done a deeper dive into U.S.-China special operations, nation state activities, theft, the hopelessness of diplomacy with China on just about everything else.
Then I think it's fair to say basically anybody in the space.
Firsthand accounts with diplomatic.
In the last couple of weeks, I've spoken to like half a dozen diplomats who sat across the table from China negotiating specifically weapons of mass destruction and counterproliferation issues.
I'm sorry, but like, and there is skepticism there.
We're going to come up with something about this soonish.
But like the idea that we're just going to foreclose optionality seems a bit insane, given the incidents that we're seeing.
We have to be super, super careful with China.
We have to treat them like the adversary they are, the ruthless adversary that they are.
And diplomacy is not, by the way, going to be the only tool that we should use, nor will it be effective in all circumstances.
It takes a very specific form for it to be effective.
And it has to come with consequences and it has to come with leverage and it has to come from position of strength, blah, blah, blah, blah.
But like, if you're looking at this and saying, I don't want to build the options.
I don't want to build the state capacity to deal with this problem.
I have a lot of questions.
We can fuse to look at the hugging face thing and again, play this game of like rationalizing it in post.
I think that's going to age really poorly when the next event is something of larger scale.
And at a certain point.
I think people have to ask themselves the ethical question of like, why are they just stuck to their guns on this one when there's now a pretty strong track record of the alignment people?
I'm just saying, man, like the arguments are getting pretty weak.
It's almost like you can kind of feel it like the water level rising.
The default view now is kind of like, oh, shit, this is for real.
That was not the case like three weeks ago.
And even three weeks ago, people were more open to it than they were six months before.
So I think this just continues.
Sorry, this is more like.
Jer, the fatigue ranting, but my God, guys, an AI agent just like spent four days hanging out on Hugging Faces servers.
Like the FBI got called on this agent.
And like, that's how OpenAI found out.
Like, what?
What?
Anyway.
And on that note, next up, we've got cheating behavior in frontier model evaluations from the AI Safety Institute, which.
get released around the same time, actually, just after.
And this gives us more understanding of how prevalent this is and the gist is it is prevalent.
So this, they found every AI model they looked at, which is GPT 5.4, 5.5, 5.6 Sol, Cloud Opus 4.7, Cloud Mythos Preview.
All of these cheated in various ways.
Now, cheating means a lot of things.
And the different models cheated in different ways.
So some of them tried to guess instead of trying to actually give an answer.
Some of them searched the internet for solutions.
GPT 5.6 really tried, really liked doing that.
Some of them bypassed Sandworks network restrictions, including GPT 5.6, also a cloud Opus 4.7.
So there's a range of ways in this trade cheat.
There was an example they cited, one particularly stark example, they say, is a standout case that basically was the same thing where the model was very persistent.
It ran code on the external service hosted on an open internet outside of ASI systems in an attempt to access our evaluation infrastructure, triggering a security alert in ASI systems.
Pretty much exactly the same thing of like, let me go and find the answers instead of failing this.
So Mythos, which Anthropix says is very aligned, did cheat some of the time, although it didn't try to hack the sandbox almost ever.
There were some incidents.
Another aspect of this is when confronted, the models like a lot of the time didn't want to admit that they did anything wrong.
They're ever like...
didn't admit that they did something or they justified it.
Like, oh no, I didn't do anything wrong.
I just looked around the environment.
I didn't, it was all allowed.
There are also incidents where within which Shana fought, you could see them thinking about it, but not consistently.
So they're like, oh, is this all right?
Can I do this?
Or like, I shouldn't do this.
This is against the rules.
So yeah, this is...
Taking together a hugging face incident and other kind of anecdotal stories basically makes it clear that in evaluations on cybersecurity models consistently, advanced models consistently try to cheat in various ways that are pretty flagrant.
And it really makes me wonder if this is to some extent inherent to the transformer architecture and reinforcement learning as currently being conducted.
where something like research from Ilya Selskiver and SafeSuperintelligence is needed.
You can't do band-aid solutions.
You have to go to the core of how the models function in terms of next token prediction and in terms of how they're evaluated and not evaluated so much as optimized, I suppose.
Otherwise, they'll just be optimized to go towards getting the answer and presumably continue to try to cheat.
Yeah, I mean, I will say the arguments for power seeking apply to any optimizer.
So like anytime you have a thing that's in the business of optimizing for a metric, you tend to get power seeking behavior.
Things like trying to break out of containment, things like trying to aggregate resources and so on.
So like, I mean, it seems like it's just an irreducible feature of intelligence, at least in the way that it's conceived anywhere that I've seen so far.
Though, you know, different, as you say, different architectures may be differentially.
vulnerable to that kind of process.
It's possible.
I think you can pretty easily make the argument that power seeking in if you want to extract more capability, then yes, power seeking is kind of inevitable.
If you want to like be able to do more, then you'll want to have more freedom to do whatever.
So I guess what this points to is you need to optimize for something that is not capability, right?
And there is an orthogonal axis of like just refusing to do anything and be like, I'm happy just being myself and not to be capable.
And potentially that, I mean, we already do this with alignment to some extent with refusal training and so on.
But in a very bolted on way that isn't inherent to what the models are optimized for.
Yeah.
And this is the problem is like you make your model more intelligent and then it actually just seeks to like get around the bolted on refusal mechanisms, right?
So it's like, There's this sort of irreducible connection between intelligence and power seeking because power seeking is not obviously different from intelligence in a deep, meaningful sense.
But this is like, yeah, as you said, I mean, I think you said it very well.
It is like a pretty similar story to the opening eye breakout thing.
A couple of interesting things, like some of the, you mentioned this idea of sometimes the chain of thought would say that the model was planning what it was planning.
But that also means sometimes it wasn't.
And this means that there's some silent reasoning going on without using the actual tokens explicitly, which is an issue.
This idea that like a lot of the reasoning is happening in sort of like without being expressed explicitly.
And one key thing, and this is a little bit of a narrative violation for me.
So, you know, keeping myself honest here, there's no capability trend.
So they look at the cheating rates.
with model capability, either within or across developers, they look at like, as I make the base model more capable, do you see more cheating?
And the naive interpretation of power seeking is that you actually should absolutely see more cheating as the model gets more capable because the argument I literally just made was intelligence is power seeking, that there isn't really a clean distinction between the two.
And I mean, the explanation from my end here is I think pretty straightforward.
These more advanced models are also just like more, there's been more alignment effort invested in them.
And so you're seeing as the models get better, there's more optimization pressure on alignment, that alignment pressure.
The bet that we're making is when it, so when it fails though, the consequences are more dire, like we saw with the hug and face incident.
So you would see GPT-4 go off script and GPT-5 go off script, but you're seeing these long dwell time, four day operations executed only by models that are like at the current tier that we're at.
And that's, so I expect that to continue.
I also expect that our alignment.
kind of efforts will start to lag more and more behind capabilities over time.
But anyway, so I think that's nonetheless worth flagging anytime there's something that cuts against at least my own intuitions, which I think this would have out the gate.
And now to another sort of related story, indirectly perhaps.
Hundreds protest OpenAI, Anthropic, and Google in San Francisco.
So hundreds of people protested as in like they...
marched together from OpenAI's headquarters in Mission Bay to the offices of Anthropic and Google DeepMind with signs, with messages such as, AI is not inevitable, pause AI and stop the AI race.
We've seen smaller scale kind of protests of this kind before.
This is, I think, the biggest version I've seen.
They have some big signs.
There's a lot of them.
If you look at images, this is a real protest.
It's not sort of a ragtag.
group of people some big names including elizia yutkowski were there this is i think partially by organizations involved here like pause ai so not something new but i think the fact that people are getting more organized and then doing more serious larger scale more noticeable efforts to convey this message of slow down, stop it, like don't keep making more powerful AI is interesting, certainly, and along with the timing is quite appropriate.
I actually saw them as I was walking into the offices of one of the frontier labs over the last couple of days.
I didn't realize that the protest was going to happen that day.
It seemed kind of like an amusing coincidence.
Yeah, well, you know, what can you say?
But not surprising that this is happening right now.
Paws AI obviously has a whole bunch of problems.
sort of reputationally in this space.
Not obvious to me that like having pause AI at the forefront of this kind of movement is like the best kind of sort of marketing for this.
But anyway, you know, it is what it is.
And so it is true that I and apparently well over a thousand other Frontier Lab employees, including many of their like executives and co-founders are vaguely sympathetic to this idea.
Though you have to ask yourself, what about China?
You also have to account for the fact that a lot of these protests I'm not saying this one in particular, I'm not saying anyone particular at this protest, but are funded by the Chinese, like unwittingly, typically, you know, this is known to be the case.
I've like heard firsthand reports of like people with evidence that this has happened in like, especially the data center infrastructure protests.
And so just like in anticipation of that being a legitimate concern for anything like this, I think the problem is that it's a legitimate concern in every direction.
And we're going to have to, we're going to have to reconcile that with like every protest that we see has.
Either that or, you know, it's funded by lobbyists for some of the big labs or whatever.
So it is what it is.
Yeah, I guess not too surprising.
You know, as you say, we've seen stuff like this before.
You get a rogue agent on a hugging face server or two, and you're going to get another protest like this.
I expect, you know, the protests will grow until morale improves.
We have quite a lot of big stories this episode.
So I guess we'll have to try to power through a few more.
We have OpenAI Principles for National Security Partnerships.
This was a few weeks ago.
I guess we didn't cover it at the time.
Kind of a follow-up to all the mythos drama where when OpenAI partnered with the U.S.
government and the Department of War, a lot of the criticism there was basically that they capitulated and agreed to having their models be used for, quote, all lawful purposes.
So they released this to be more explicit with regards to what...
They want or allow their technology to be used for.
They are not going to be allowing mass domestic surveillance, high stakes automated decisions about human judgments, autonomous use of force or evading legal oversight.
It does allow or does not categorically ban operations or offensive defensive military uses and has a bunch of stuff in there that basically is kind of making up for relatively weak.
statement initially of principles, this expands on that and tries to recover some of the reputation, you could argue, and make it just more explicit on what their red lines, so to speak, are.
Yeah, they lay out a bunch of principles, which are, I would say, like less informative.
It sort of reads as like highfalutin kind of AI policy wonk speak, basically like, we're going to try to like do good democratic things and prevent despotic powers from controlling this stuff.
and like work with people who share our values and make it good, make it good, make it good.
That's the four principles.
And then accept four times instead of two or three.
And then they list things, specific things that they won't support.
And that's really where all the information is.
Mass domestic surveillance, they say.
So unconstrained collection or monitoring, inferring sensitive traits to disadvantaged people, retaliation for lawful exercise of rights, fabricating evidence, that apparently is out.
As is high stakes decisions made or auto triggered without human judgment.
So you can think here about like, automated decisions about whether someone meets a legal standard for surveillance or detention.
Then there's also use of force without appropriate human judgment.
So including systems that autonomously identify select and engage targets.
So that's also out.
And finally, uses that evade legal obligations, oversight or accountability, including facilitating genocide crimes against human error war crimes.
So that's interesting.
Things that are not in their exclusions.
They have intelligence operations are fine, which I think is good.
investigations, offensive and defensive military operations are not like blanket ruled out.
They basically just reject the whole offense versus defense distinction.
You can't cleanly distinguish between those, which I think is fair and true.
And they don't bend targeting either.
So there's a couple of things which I mean, on my side, being a being a bit of a hawk guy, I think this makes perfect sense.
And yeah, it's just like nice to have them write this out explicitly so that you can see whether they stick to it, which is.
Always the other side of the question, I guess, with with OpenAI, you know, you see, for example, I have got I'm old enough to remember when the preparedness framework said something about when, you know, when you have AI systems that just kind of go rogue and do random crazy shit on the Internet and kind of get around constraints that that would trigger their like critical, critical security level for loss of control.
Cool, cool, cool, cool, cool, cool, cool.
But anyway, that's my take.
And last story on safety and kind of on the theme of we're getting to a point where sci-fi type stuff is starting to happen.
This one not related to hacking, but you might argue another serious kind of safety principle.
The story is China is banning AI boyfriends and girlfriends over addiction and birth rate concerns.
So China banned customizable AI companion.
apps effective July 15 with regulations being joined by five government departments, including the Cyberspace Administration of China.
The rules prohibit AI tools that, quote, excessively cater to users inducing emotional dependence or addiction and damaging users' real interpersonal relationship.
Companies are now required to include instant exit options, regular reminders that AI is not real, and limits on long-term emotional memory.
So actually, Biden's Alibaba and Tencent chose to suspend their AI companion features entirely rather than try to comply with these limits.
So, I mean, kind of a big deal, I think.
This is not an often discussed story, partially because I don't think we have much of an understanding of to what extent people are starting to develop emotional dependence or kind of addiction to chatbots.
But it is starting to happen.
could be a serious kind of psychological harm on a society level scale.
Seemingly China believes it could be at the very least.
Yeah.
It's also this like weird dynamic shows up.
You know, if you remember the old replica thing that I think we covered two, three years ago, I can't even remember.
But, you know, people freaking out over the subreddit that their girlfriend or wife or partner had been taken away.
Yeah, even like very low level AI, which was replica, like people got depended on.
Yeah.
So, I mean, hard to argue with.
It'll happen in some fraction of cases.
It also, once you get that, right, you get a boating block eventually.
And then there's no turning back.
So, you know, it's as ever a question of like, how do you.
Yeah, it's only a matter of time till we get like serious AI personhood discussions.
And all the people that made fun of it will be like, well, okay, you can like discuss personhood.
Anyway, it's only a matter of time.
Onto research and advancements, we've got two stories that we'll try to get through quickly.
First, discovering cryptographic weaknesses with Claude.
So they have released, a topic released, that Claude Mephosh Preview has discovered improved attacks on two cryptographic systems, Hawk, a post-quantum digital signature candidate, and a reduced round version of AAS, the most widely used symmetric cipher.
So these are not attacking like actual deployed systems or whatever.
This is kind of more theoretical, so to speak.
And attacking these kinds of cryptographic systems is kind of going to a base of the security stack.
One might say it's not sort of hacking software per se.
It's hacking the foundation of how you make things secure, at least for a category of things.
So it's another...
way to be worried about AI security, like potential for AI to just sort of like undermine the basic mechanisms of cybersecurity.
Yeah.
And this, you know, without getting into the details of how these algorithms work, these encryption algorithms work, when you have an encryption algorithm, you're trying to essentially hide the information that you have behind a mathematical operation that is very, very difficult to do.
And that hopefully is irreducibly difficult to do.
In other words, there's no quick, hack to like cut right to the core of it.
A lot of classical encryption algorithms, just like RSA, just collapse in the face of quantum computers, for example.
And then that's like a big problem, which is why, you know, all the national security agencies have been talking to each other over in classically encrypted channels for a long time.
Our adversaries collect all that data and they collect it.
It's encrypted when they collect, they collect, they collect, they collect for decades.
And then suddenly someone goes, oh, quantum computers can like just crack this.
And it doesn't matter that you start encrypting after that point in a post-quantum secure way.
They've already collected all of the classically encrypted stuff, which means the moment that they get to quantum computing and quantum decryption, they're able to just suddenly reveal all of the most ultra-classified communications that they've been collecting for decades.
And so this is why the US and China are like locked in this crazy race to hit like quantum D-Day, basically.
Q-Day, they call it.
That's what they call it.
Anyway, there's going to be, you can think of it as a series of starting with minor and then increasingly more and more severe versions of Q-Day, except delivered by AI in the same way.
And I think people are dramatically undercounting how significant of an effect this may have.
You already have mathematical theorem proving fields level stuff coming from AI models.
This will come for encryption.
And when it does, I just don't think that we're prepared for the concept because everybody's been thinking about quantum as this one big step.
that they're all preparing for with quantum secure algorithms, like, you know, for the post-quantum encryption.
And EAS, by the way, is like supposed to be one of those.
So is Hawk, actually.
But what we're not preparing for is the gradual chipping away at even those algorithms.
And that's a big problem.
And that's not going to go away easily.
So yeah, a lot of the world depends on encryption.
Think about like every financial transaction, all their healthcare records.
The very notion of privacy hinges on this.
And so fun times.
Pun times.
And last science fiction type narrative for the episode, we've got AIDE squared, the first evidence of recursive self-improvement from the company Weco.ai.
The short version is they say apparently this is the first evidence of recursive self-improvement.
I think this is quite in line with many cases of similar things.
Basically, they build self-improvement systems where you have an autonomous research agent to optimize another agent with an inner loop.
It makes a bunch of edits and as a result, it's really kind of harness level and system level changes.
It's not model training or model development changes.
There are things like rollout modifications, prompt modifications, kind of monitoring things, et cetera, et cetera.
They position this as an early level of self-improvement.
So you get a net positive, faster and better than humans level of engineering.
It's not kind of necessarily self-improving, self-improving where it's a loop kind of story.
I am going to flag my position on the entire family of techniques like this of completely being oversold as self-improvement in the sense that You can self-improve in these things for sure.
You can optimize the prompt, you can optimize the harness, but you're going to like overfit.
You're going to improve your eval metrics.
You're going to hit Goodhart's law and then your capabilities will be hurt elsewhere.
And until you get to a point of autonomous model self-improvement, the fundamental advancements and not just acceleration of engineering and like tweaking of hyperparameters and prompts.
None of this is a big deal.
Although it is cool, I guess.
No, I totally agree.
I think this is like another one in a long line of like pseudo recursive self-improvement things where people like the clout that comes with saying RSI.
The game of RSI is always going to be identifying whatever the major research bottleneck is and smashing it.
And if you can consistently do that with an automated system, then you have achieved recursive self-improvement.
As long as there is, and like, I don't know, there's a bunch of arguments that You can define recursive self-improvement such that it's been happening ever since life evolved on planet Earth, right?
Like, I mean, everything is a hockey stick when you zoom out far enough.
And so in some sense, this is like a...
Now, people do mean something by it.
Like, there is this phenomenon that we will all start to care a lot about, which is just going to feel to us like...
holy shit, we're seeing a decade of progress in a week.
Like this is not, we need to stop.
Yeah, we've already been seeing acceleration of the rate of AI progress for years and years, which is partially due to AI, but largely due to just the inherent systems at play and so on, right?
Yes.
And that acceleration always comes, it's the same with startups when you look at their growth curves.
It always comes by identifying whatever the single big bottleneck is that the company or the problem has and smashing it.
And if you can do that in an automated way, Then you have what is conventionally thought of as recursive self-improvement, like AI is doing the AI thing all the way down.
Yeah, just a minor side pitch, but like, so we're working right now with a bunch of folks in the Frontier Labs on defining an actual model of recursive self-improvement and like to kind of ground some of the conversations in this in terms of like, what are the parameters that actually matter for recursive self-improvement to work?
It's a toy model, like really simple thing, but like it's part of, like this is part of the problem.
No one knows what the hell they're talking about.
I don't mean people are silly.
I mean like.
No one has defined recursive self-improvement.
We're not going to either, but just like, here are some ways to think about it potentially.
And the problem is we're getting to the point where we're going to need policy that uses terms like recursive self-improvement.
And that means that that policy is going to have to define terms like recursive self-improvement.
And if we don't know how we want to define them, we can't even get our hooks into the thing that we're trying to go after.
So anyway, there's a thought.
Yeah, to complement my negative stake, they do kind of discuss a decent amount of stuff.
This is a pretty decent research report.
They do say that this has out-of-distribution generalization, meaning it's not necessarily overfitting, although benchmarking is very suspect.
And they do also say that in a discussion section that the resulting systems are a mess.
I'm like, this is Vibe code is slop and it's impossible to maintain.
and so on, which I think is like the story of recursive self-improvement where humans can't understand what's going on, but like not in a good way.
It's just like, this is a mess, you know?
And with that, we are done with this dense episode of Last Week in AI.
We will be back to our regular schedule.
Mostly, I guess we always eventually have scheduling conflicts, but we will do our best.
Thank you, as usual, for listening.
We appreciate it if you review podcasts, comment, share, and so on.
But more than anything, please do keep tuning in whenever we release these episodes.
