# AI Geopolitics, Enterprise Lock-In, and Safety Risks

**Podcast:** Last Week in AI
**Published:** 2026-03-16

## Transcript

Last week in AI, I would like to thank ODSC AI for being a sponsor.
ODSC is one of the longest running and largest communities focused on applied data science and AI.
It started over a decade ago with a simple idea bring practitioners together to learn from people actually building and deploying models in the real world, not just talking theory.
On April 28th through the 30th, you can experience that yourself at ODSC East 2026, taking place in Boston and virtually.
There will be thousands of hybrid attendees ranging from data scientists, ML engineers, AI researchers, and technical leaders.
You can attend over 300 sessions covering LLMs, Gen AI, computer vision, NLP, data engineering, and more.
You can also go to hands-on training with workshops and boot camps taught by experts from companies like OpenAI, Hugging Face, NVIDIA, and other top companies and universities.
And of course, there'll be a massive expo and networking opportunities, great for startups, hiring managers, and AI tool builders.
It's one of the best ways for AI practitioners and teams to stay ahead of a field, learn from the best and connect with a community.
Go to ODSC.ai slash East and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.
That's ODSC.ai slash East and use code LWAI to get an extra 15% off on the number one AI builders and training conference.
We'd like to thank Fox for sponsoring Last Week in AI.
Fox's leading intelligent content management platform, enabling organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI.
To unlock the power of AI, you need to get your content to your LLMs and agents.
Your business isn't the sum of internet knowledge.
Your business lives in your content.
So you don't just want to bolt on AI to your existing processes.
To become an AI first company isn't just about automating what you already do, it's about reimagining what's possible.
With Box AI, you can truly leverage the latest breakthroughs in AI to automate document processing and workflows, extract insights from content, build custom AI agents to work on assignments and more.
And most importantly, Box AI works with all the major leading AI model providers, so OpenAI, Anthropic, Google, XAI, and others, so you can be sure you can use the latest AI models with your content.
Boxai will give you the content layer that gives AI the context it needs while giving your teams the flexibility they need to test and leverage various models for different use cases.
So go to Box.com slash AI podcast where you can hear us chat about what's going on with AI.
As usual, in this episode, we will summarize and discuss some of last week's most interesting AI news.
You can also check out our last week in AI newsletter at lastweekin.ai for even more news.
I'm one of your regular hosts, Andre Karenkov.
I studied AI in grad school and now work at the startup AstroCade.
And I'm your other co-host, uh Jeremy Harris.
I'm at Gladstone AI doing AI national security stuff.
And yeah, we so last episode, right?
I think we we ended up doing just news stories.
We didn't do any research, any papers because of just how much there was.
And so this week, it's kind of the opposite situation where we're just like we got a shit ton of papers to cover.
Um so we'll see what we can do in in time.
But this is gonna be an interesting challenge, yeah.
We're really kind of vacillating from one mode to the other.
Yeah, it's a week where there was not a ton to say, no new models on like you know, previous weeks, uh, not many business stories.
So you're probably gonna get through those sections pretty quick and then get to talk to about like six papers or something.
It's gonna be a lot.
Before we get there, just want to call out, did see a new review on Apple Podcasts, super solid and informative, a super concise review, which is pretty cool to see.
So I appreciate that.
We do our best to keep this super solid and informative, and uh do appreciate all the feedback and also the comments on YouTube, which I keep an eye on.
I will say it does it really does help psychologically because it's like it is a lot of work to prep for these things.
And so seeing that it's appreciated feels really good.
And and I love that sense of community too on the uh on the the um Apple side and also on the YouTube side.
Like it just seems like there's a lot of folks chatting.
And anyway, it's it's it's great to see.
So really appreciate it.
And kicking off with Tools and Apps, first story is perplexities, personal computer turns your spare Mac into an AI agent.
So they have announced this new AI tool called Personal Computer which is basically OpenClaw by Perplexity.
So it's gonna use your Mac Mini which is famously what OpenClaw runs on.
You get a Mac Mini, you uh set up OpenClaw and it kind of like lives on that server effectively.
So Perplexity is positioning this as the safe alternative you know the actually kind of more adult mature version that you can use as your personal assistant.
You can provide it full access to files and apps and uh you know there's a nice looking UI to use it.
So interesting that they kind of I don't know if they pivoted or uh rushed into this right now it's not available to the public you can join a wait list for early access.
There's no specified launch date.
So it really seems to be like, oh, this is probably a good idea.
Let's do it.
And we'll see when it actually comes out.
Yeah and and this is obviously hitting on the head the big security issue that uh open claw has had.
Everybody's been talking about it.
You're basically giving this thing unlimited access to everything on your machine.
It runs on a cloud, blah, blah, blah.
Well, this is going to be a local version of that, as you said so so the idea here is to fundamentally shift the security picture so that, as you as you say, it's sort of the mature grown-up adult way to do this.
There's a lot to the security picture.
I think that'll this will be a really interesting space for startups to play in.
Having open claw run in ways that are verifiably secure in ways that are updated very quickly to account for new capabilities and trends in AI capabilities, because you know, new security issues, as we've seen just seem to be popping up all the time, almost as fast as you can go, which makes a really good case for a subscription service.
I mean, this is like the classic, uh, the classic reason that you would subscribe to something.
So, yeah, uh really interesting.
And as you say, a big strategic move uh by perplexity.
Not clear they have the option not to do this, because plain old search is not gonna be the way of the future for for much longer.
And next up, a new update to Cloud Code.
They have launched this feature code review, and that is what it sounds like.
You can now have it review code on GitHub, and there's this thing called pull requests, which as a software engineer you do, so it can automatically review those and uh provide actionable feedback.
Saw some kind of funny discussions of it on Twitter where people are saying it's like costing 15 to 25 dollars per code review, and also that this is sorely needed to deal with a glut of code being pushed out there by people who use cloud code to write the code in the first place.
This is joining a pretty busy space.
There's multiple companies that are entirely working on this, like Code Rabbit with uh AI code review automation.
Yeah, and uh yeah, not surprising.
Uh, this is in some ways more of a nice utility because you can already use it for this purpose.
And in fact, at my company, we already do use it for this purpose without having any sort of fancy UI or automation behind it.
Yeah, this is a pretty interesting move.
It is positioned as a safety tool, right?
So it's like and security, you know, catching bugs, um, looking at AI generated code and so on, but it's quite interesting that this is also a growth flywheel uh in its own way, right?
The more clo code that Claude Code generates, the more PRs are gonna pile up, and the more companies need code review.
And so uh anthropic's kind of like closing that loop in a very interesting way, all under the same roof.
They are also just basically looking at a multi-agent architecture by design, right?
So the the core philosophy here is get a bunch of agents to look at the code base from different angles and and kind of a final agent that aggregates ranks and findings, and they're exclusively focusing here on like logical errors rather than style, which suggests that they've they've learned from the graveyard of linting tools that developers have ignored or developed um uh over the years.
And so uh, yeah, I mean, this is quite interesting.
Uh the the pricing is is interesting as well.
So each review, each PR review is estimated to be between 15 and 25 bucks.
So if you've got, you know, a big company with hundreds of engineers, each shipping, you know, 10 plus PRs a day, that's you know, it might compound fast.
It's quite interesting.
The cost, of course, of doing a PR review, if you look at the hourly rates of uh developers are certainly going to be in that range.
So it's kind of interesting.
You know, we've talked before about the the cost of getting AI to solve certain problems, sort of flirting with the human cost already, and that'll obviously collapse too as AI gets cheaper, and Moore's law and Huang's law start to play into this.
So pretty interesting.
This is um also anthropic doubling down on the enterprise side following the Department of War lawsuit, right?
So doubling, tripling down in that direction that's working maybe to ensure that they have a not a fallback.
I mean, that's the wrong frame because they are incredibly successful in what they're doing here, but just to kind of make sure that they're they're boosting that part of the business.
Yeah, and I think worth mentioning kind of a broader context of they've been just pushing a ton of cloud code updates in recent months.
Uh they also, I think this week just launched was like buy the way feature where you can have a little side chat as you're uh having Claude do something.
So it's very clear that Cloud Code is now one of our main products, really generating a ton of revenue, that's seeing crazy adoption, and they're really building it out in a way that wasn't the case last year.
And next, uh sort of related story, cursor is rolling out a new kind of agencoding tool, although that might be overstating it.
They launched a new tool called automations, and that is what it sounds like.
You can launch coding agents based on triggers like code-based changes, Slack messages, or timers.
And by the way, actually, I think Cloud Code also has this where you can schedule tasks and have it do it regularly.
Yeah, yeah, yeah.
So uh that's pretty much all there is to it.
You can now uh have agents be triggered by various things, not just you directly prompting it.
You could probably connect it to be like, oh, there's a new pull request, go have it do the code review.
It fits into this broader trend of having agents sort of live out there in a server or a cloud, and uh you can kind of ping them and have them do stuff, as in OpenCloud, for instance.
Uh, although in this case, the difference is it's automated, so they're just running 24-7 and doing stuff uh whenever.
Yeah, I mean, uh, you can think of this as as kind of like a shift from like humans as the orchestrators to or generally like it's shifting from agentic or orchestration to an infrastructure play.
So, right now, most agentic coding tools have a human that who serves as like the dispatcher, right?
So you you write a prompt and then the agent runs and then you you review.
And the whole idea here is you're kind of reframing the human's role in this.
So instead of just like initiating what you're doing is you're you're being, as they put it, called in at the right points in the conveyor belt, right?
So you're kind of just your attention is being drawn to specific moments in the workflow where there's a human inject that's really required, which makes this a you know pretty different mental model from what software development has looked like, even in the AI augmented age.
So, in that sense, like this is part of cursor being positioned right now is you know, two billion dollars in annual revenue, doubling in three months.
You know, they're they've got uh a huge fraction of the generative AI market for coding.
They have enough trust that they can start to reshape workflows and not just augment them.
And that's what they're leaning on here in part.
I mean, we're gonna see more of this, and as you say, like anthropic's already knocking on the same door, but this is kind of a very opinionated stance on what the next level of abstraction is going to be and sort of seeing human attention as this thing that you're very, very carefully dialing in.
Um again, less orchestration, more sort of AI as infrastructure.
Right.
The way they position it is always on agents, which I think is a nice term for it.
Again, this is similar to OpenClaw and now a perplexity's computer.
The agent is out there and you don't need to be at your computer, you know, talking to it in like chat GPE.com.
There's various ways to get it to do work for you up there in the background, and you can kind of join in and see what it's up to whenever.
And hopefully it's not up to deleting your emails or wasting your money, as can happen with OpenClaw.
Next up, ChatGPT can now create interactive visuals to help you understand math and science concepts.
Uh it can create interactable real-time visuals.
Apparently there's 70 topics that it has by normal square, Charles's Law, Ohm's Law, and other things like that.
So uh sounds like maybe it's kind of a pre-built set of visualizations, presumably for things that people often ask it about.
Uh and yeah, it's like a little mini app where you you have ways to visualize the actual logic behind equations and so on.
And related to that, actually, on Tropics, uh, Claude can respond with charts, diagrams, and other visuals now.
So the visuals here are gonna be inserted in line within the chat, and it appears to be uh more general.
So it is a little mini program that it's creating for you to be able to interact with.
Uh the example they give is you can ask for create an interactive version of a periodic table, and it does that.
Presumably, this is not really different from artifacts, which uh Claude has had for a long time, where it can code up a little app, but the difference is it's actually there inside the chat directly in response to your question.
This is anthropic finally covering down on a gap.
You know, if you've used Claude, you've seen this, right?
It's like it doesn't do audio image and video models, and you know, it's like mostly just not the model that or the um company that you turn to for those sorts of things.
And this is them trying to kind of write that wrong.
The timing is interesting, obviously.
You mentioned this, you know, days after OpenAI updated Chat GPT for the science visuals and and math concepts, though that is also kind of a more specific use case.
This really is the broader issue of anthropic not having had much of a strategy for kind of images and videos.
And the other piece here is they are differentiating by making their visuals conversational.
So you actually like see them update as you have your discussion rather than just having a static output.
So, you know, uh all that is is quite interesting and uh just trying to broaden the market as you know, anthropic hasn't had to, because typically you know, being more business to business like B2B versus B2C, um business to consumer, Anthropic has been able to get away with you know focusing on code, focusing on on reasoning, that sort of thing almost exclusively.
Now they're you know kind of broadening out a bit.
And um, I don't I don't know if this is uh an explicit play to go more consumer.
This is gonna have B2B applications as well, obviously, but um yeah, interesting in that context too.
And uh now actually a different section than usual.
I figured we'd throw in projects and open source.
We have one story here, and it actually is related to tools, in that you know, we have a new model out there introducing Namatron free super and open hybrid Mama Transformer MOE for genetic reasoning.
Uh, this was announced by Nvidia.
Nematron free super has a hundred and twenty billion total parameters with 12 billion active parameters per inference.
It has apparently uh one million token context window.
Although I I would do want to mention that when we say one million token context window, as we have GPD 5.4 and I think Sonnet 4.6, you know, on paper it's uh about uh 1 million in practice as you get into you know the upper end, maybe the model starts being very stupid.
The model compares favorably to other open source models, although probably not all of them.
The charts that they provide that just compare it to GPT OSS 120 and Quen 3.5 uh 122.
So in that uh size class, it is on the benchmarks doing well, and the throughput is way, way higher.
So that's one of the cool things with uh both an MOE model and a Mambo model.
The performance is comparable, but it can be run locally very, very quickly.
Yeah, they're also this so they're training natively in four-bit arithmetic.
So usually what happens is you'll see like FP16 or like you know, you know, BF16 or whatever, where like these higher resolution numerical representations of the weights and biases in the weights in the model.
And then what you do is after the fact, you will then quantize the model.
You'll essentially create a lower dimensional or a low resolution representation of all the numbers.
And you gotta hope that it it operates about as well.
Uh, what they're doing here is actually like training natively in four-bit arithmetic.
So just four-bit quantization, which is helpful because what tends to happen when you try to go from a high resolution numerical representation to a lower one in like in one quantization step, is that like the problem is that the model was trained to leverage the resolution that it can enjoy, right?
With whatever 16-bit resolution or whatever.
When you flip it to six to four-bit, well, maybe the model would have made different choices actually during training uh in terms of the values of some of the weights if it had known that it was gonna be used at 4-bit arithmetic representation.
And so they're actually pre-training fully in 4-bit, so it's it's quantized already, coming out of the box, if you will.
And what this does is it means that this model is optimized for black wall GPUs specifically.
So they are actually kind of interleaving their model design and the hardware very, very closely.
It's totally rational from their standpoint, but it's worth thinking about if if you're thinking about portability, this model will not work as nicely on non-black weld GPUs, for example, which is sort of the you know, a kind of hardware lock-in play.
It's open source, right?
This is NVIDIA saying it like open source, but but you can use it only on, but not only on, but the performance will be best on their hardware.
That's a really interesting play from NVIDIA.
There's also this idea of latent MOE that they're doing.
It's a bit of a sleeper feature, but you're basically compressing, so you have your token embeddings, you're going to compress them to a lower rank space.
Essentially, you think of it as a lower dimensional space before you do the MOE thing.
So the MOE thing is you take your token and then you pick any number of a you know bunch of expert sub-models to route it to, right?
Now, in this case, what you're doing is you're compressing the representation of your token, and then you're routing it to these models.
And what this means functionally is you're consuming less compute.
The token it consists of fewer numbers that have to be crunched on by those experts.
So you actually, if this is cost less compute, it means you can, for the same amount of compute, uh consult about four times as many experts for the same computational cost.
So this is a big advantage because oftentimes you know you you have too few experts chewing on your data.
Ideally you'd want as, well, to some extent as many as possible.
There are orchestration issues that come with that.
But yeah, so it is quite interesting.
It means you know you have finer specialization.
For example, you have like distinct experts, uh routing paths for like Python versus SQL logic within a single agent turn.
So that's pretty meaningful, especially when you're you're thinking about complex tool calling agents.
So quite interesting.
Uh I think the one million token context window probably deserves a fair bit of scrutiny when you're you know claiming a large context window, like actually using it faithfully, always a different thing, right?
We've we've seen Mamba architectures with known weaknesses in like multi-hop retrieval over long distances.
So you know, and the benchmarks that they're citing here are they're pretty narrow, you know, like Pinchbench, for example, is new enough that it hasn't been stress tested by the community more broadly.
But but anyway, so so this is something to look at.
You will see what the sniff test is uh as things evolve.
But an interesting model for sure in NVIDIA, continuing with this play, right?
I mean, they're gonna they're gonna continue beating that that open source drum because the reality is they dominate the open source uh market way more than they they already dominate the closed uh source model market, but they really dominate the uh open source market and want to get their hooks in with again these models that are most performative on NVIDIA machines.
That's one of the big take homes from this paper.
Right.
And and by the way, the actual news here is that they released the weights for Nematron 3 super.
The paper was released in late December.
I believe we covered it uh a couple months ago and discussed Nematron 3 in more detail.
Uh, one of the other interesting details is they're very much doubling down on this hybrid architecture of Transformer with Mamba, which uh is one of the ways in which it can be made efficient, uh, keeping compute kind of constrained as you scale up a number of tokens, which could mean that it's actually pretty good at long context reasoning.
But it also seems pretty apparent that they are not at the point where they have all these post-training tricks of reinforcement learning.
If you look at the paper, they have a pretty large pre-training data set, but they don't discuss.
Uh they do have some post-training, but I just don't think their infrastructure is as mature.
And one last thing to say latent MOE, they released a paper on that uh back towards the end of January.
Uh, I don't think we actually covered it at the time.
And it is in a way related to a hybrid approach.
They position it as revisiting MOE design, uh mixture of experts design from a hardware software co-design perspective to make it optimal with respect to inference cost.
So, yeah, I think interesting from Nvidia in that they are kind of carving out a different space in how these open source models are.
They aren't the most performant, but you can presumably deploy them on a single GPU and get really uh nice kind of performance, which does work well with what NVIDIA sells, obviously.
And onto applications in business, sticking with NVIDIA.
Nvidia halts H200 production as China backs Huawei chips.
So it is stopping production of its H200 AI chips for China because of uh China, yeah, as as the headline says, kind of uh not really being welcoming to their companies adopting these chips.
This is the latest generation of Nvidia chips.
So these are top end chips, not the H20, which is kind of the low end that was meant to comply with regulations.
The US has approved limited exports of these chips to China, but Chinese customs block their entry and uh suggested that domestic alternatives should be prioritized.
So, man, Nvidia is really you know shifting waters.
Uh did it seem pretty clear what's going on.
Now it's like, oh, we can export them or we can't export them, and now we don't want them, but the black market is probably good, so we'll actually still have them be sold there.
Uh I don't know.
Yeah, no, I mean this is one of those situations where I'll just keep like my personal opinion here is that this is a kind of two wrongs make a right situation.
So I don't think that we should ever have been approving the export of H200s to China.
That just like that seems insane.
For context, right?
Just as a reminder, the numbers here are super confusing.
So the first there was the H100 GPU, right?
This is the like the um GPU that would have been used to train models in the like you know 2024, 25 era, right?
Okay, so that's the H100.
Then NVIDIA came up with China specific variants of it, the H800 and the H20.
Those were shipped to China in large volumes.
There's all kinds of interesting arguments about how those chips in some ways were actually more performant than the H100, which is a huge issue, but whatever.
Then there was the H200, which, despite the fact that it looks a lot like H20, it's just missing a zero, was not at all intended for the Chinese market originally.
It is strictly an improvement on the H100, uh, I believe using the same manufacturing process at TSMC.
So as a reminder, NVIDIA designs the chips, they ship their designs over to Taiwan's semiconductor manufacturing company, TSMC in Taiwan, and they actually fabricate the chips, right?
Using their exquisite four nanometer process or five nanometer process for depending on the GPU.
Okay, so in this case, the Trump administration said, hey, guess what?
We're gonna over like overturn all the Biden administration stuff and say, okay, yeah, the H200, so this fully fledged NVIDIA chip, we're gonna allow it to ship to China.
So the expectation was that okay, great, NVIDIA is gonna make a crap ton of money doing this.
It'll accelerate massively Chinese AI development, which a lot of people don't think is a good idea.
But NVIDIA's like Fantastic will make some money.
Uh problem is, NVIDIA then ramps up production, right?
So they go to TSMC and they say, hey guys, we want to buy out a huge chunk of your capacity at your your, I guess your four nanometer node here to build these H200 chips.
And that's really expensive.
There TSMC has a finite capacity.
There's more demand for TSMC fabrication capacity than there is supply, right?
So getting allocation on TSMC's nodes is really, really hard.
So NVIDIA does that, and then suddenly China goes, eh, you know what?
Um maybe we do, maybe we don't want these uh these NVIDIA chips.
So on the one hand, after Trump approved these limited exports to China in January, China came back and said, well, you know, Chinese companies they can purchase the H200, but they should consider local chips first.
And now they're basically coming back and saying, nah, like we're not gonna do this at all.
And as of now, no H200 chips have been sold to Chinese customers, at least according to US officials.
And so some Chinese commentators apparently have been saying that China's decision to curb these H200 imports is due to the US intercepting over almost 2 million barrels of Venezuelan oil destined for China in late December.
So China is actually viewing their blocking of H200 exports to China as a penalty to the United States, which is hilarious because this is all ass backwards.
Like really, the national security imperative would be to prevent these H200s from going to Chinese hands.
The Chinese have way more to gain by taking these chips than to lose.
But the way they're framing this up for themselves is well, we'll increase demand in the domestic market, blah, blah, blah.
There are all kinds of reasons why, at least I mean, I think that's a completely implausible position.
In any case, so now we have a situation where there's like a ton of anyway.
I'm I'm gonna I'm gonna park it there.
There's a whole kind of um story here as well on the memory side of this.
So the H100 and H200, one of the big reasons they're so powerful than the Huawei Ascend chips, for example, is that they use HBM that's a kind of high bandwidth memory that is made by SK Heinex, the South Korean company, and is much more advanced than what China has.
Their closest competitor to SK Heinex is CXMT, but they're multiple generations behind, and there's been an attempt to kind of puff up CXMT in this article.
They're like, oh, our local solution to SK Heinex or Samsung CXMT is really impressive and blah, blah.
The numbers just don't pan out.
They're basically two or three generations behind here, not close enough to be relevant.
Uh, but this is part of what China's trying to do is make it seem like you know, they're a puffer fish.
Like we can, you know, we have CXMT, we have Huawei, we have uh SMIC and so on, and and there you go.
Everybody's everybody's like you know, stumping their chests or whatever.
Right.
And I think we've discussed previously how there's a massive black market, the chips kind of end up there, even though they perhaps shouldn't.
And the news is that they are reallocating manufacturing intended for the Chinese market.
But you know, uh what does that mean?
You know, oh yeah, sorry, I'm gonna get on my soapbox for a second here.
Um, I'm old enough to remember when people sorry, I'm getting all hockey in my couple here.
I'm old enough to remember when people like me, and and much, much smarter, all much smarter, we're saying, look, there is a finite amount of capacity at TSMC and at NVIDIA.
Any chip that you're sending to the Chinese market is a chip that is not being received by an American company directly trying to compete with China.
That is just true, and that's what the story shows.
Literally, NVIDIA is deciding, okay, we can't ship these to China.
Guess we're gonna make more chips for the American and Western markets.
Like, that's literally the calculation.
So, as far as I can tell, at least, what this suggests is that that argument was always true.
It was something that Nvidia has claimed was false, uh, that actually, you know, the it's a it's a win-win thing, and like you know, giving us access to the Chinese market is blah blah blah.
Anyway, this like it could not be clear, at least to me from the story, that that's in fact exactly what what you would expect in a in a uh sort of supply constrained environment.
And and by the way, this is all according to people familiar with the matter reported by Financial Times.
So it's not like NVIDIA made a big announcement about this, or anything it's uh you know, some internal company stuff going on that appears to be true, it's just gonna make sense.
Next up, an update on X AI.
Another XAI co-founder has left, and another one says he is leaving.
So this is now at the point where only two of the 11 co-founders of XAI are gonna be left at XAI, which I guess is now SpaceX, technically.
So yeah, you know, the people leaving are Jiang Dai and Guadong Zhang.
And these are some heavy hitter roles.
Uh they were leading the initiatives for Gra Code, at least one of them was, and I think the other one was working on Grock Imagine.
Jiang Dai previously worked at Google Brain.
His papers include Excelnet, Gemini, Transformer Excel, Gemini 1.5, so very high pedigree.
Guadong Zhang was also previously a researcher.
And uh not to drag on them, but to the two remaining co-founders do not have much of a research pedigree.
They're previously at Tesla as a technical manager, and at X Twitter as a senior director of software engineering.
So the talent leaving XAI and the leadership talent in particular, you know, these rare high demand real AI experts who have been deeply involved in things like developing Transformers and Gemini are leaving.
And Elon Musk uh jumped on uh some uh story about the saying that XAI was built the wrong way the first time around that is now being rebuilt.
So not sure what that means, but uh XAI certainly seems to be being rebuilt from like the ground up.
Yeah, I mean, it's um not a an uncommon thing, obviously, for Frontier labs to to lose their founding members in droves.
This fast is unusual.
And I will say the exception of that, of course, is anthropic, which puzzlingly has managed to retain all, you know, I forget eight, nine founding members, which is quite remarkable.
The flip side is, I mean, man, is continuity of expertise hard to maintain, right?
If you start to lose the founding minds like this, uh like research directions should be planned out over, you know, a year, a year and a half, two years, like, yeah, that's the kind of thing, but this is this is gonna cause problems.
Elon's approach on this has been to post what you could interpret as an apology on X, acknowledge that like a lot of talented candidates were wrongly declined interviews, and he says he's personally gonna go back through hiring history to fix that.
With the frame he's using is as you say, yeah, XAI was not built right the first time around.
It's kind of like re-founding the company, which is consistent with, you know, any time there is a change in executive leadership.
Traditionally, you know, when when you come in as a new CEO for a company, you have to re found the company, right?
That's that's what you do.
That's what uh that's what Satya did at Microsoft.
It's what um Andy Jassy has had to do at Amazon and so on.
The acquisition by SpaceX gives him that out where this is like a refounding of the company.
It's all coming at the same time.
Given the palace intrigue at XAI, given these challenges, you know, we have to ask ourselves like at what point does XAI start to risk existentially, like not being able to draw in the best talent, which is that is the existential question for a Frontier Lab.
If you're gonna compete at the Frontier, you need the best.
Their big solution to this is the SpaceX acquisition, right?
It's data centers in space will be the future.
It'll be the only way to get, you know, the next 100 gigawatts.
And and the fact that they're leaning on that story so hard is the pitch, right?
If you believe that infrastructure story, then you want to get in on the ground floor at XAI and not other labs, regardless of the palace intrigue.
That right now, I think is kind of the best card they hold.
It's actually it's been debated extensively.
I think it actually holds water quite a bit, surprisingly, the idea of data centers in space.
When is a separate story?
And ASI maybe reached well before we have data centers in space.
In fact, if you talk to folks at OpenAI and Anthropic, that certainly is where they tend to orient.
So all this is kind of like what is your thesis about where AI is gonna go?
That's gonna shape which lab you pick.
Right now, XAI's infrastructure play, given all the conflict there and and and the churn of employees, I think that's their best card right now from a recruitment standpoint.
And it's certainly not nothing, but people choose their thesis.
Right.
Also worth noting, uh, just as this was happening, uh Jason Ginsberg and Andrew Milich from Cursor, where they were leading product engineering, are joining XAI.
So there's some new talent coming in.
But you know, visa are not qualitatively the same.
If you work at Cursor, you're not training a foundation model.
You're not like trying to compete with Claude at Chat GPT, you're you're building something on top of that as an application, uh, which could is something like rock code, for instance, but isn't a product.
On the expertise side, it's looking rough.
Like it's not, it does require, as far as it seems to me, some kind of secret dark arts magic knowledge that isn't so public and isn't something you necessarily have unless you have experience for training models for knowing how to do pre-training, mid-training, post-training, how to get that all optimized.
So at some point, you know, you have a data center with a gajillion GPUs, but uh is XAI actually gonna benefit from it that much if they don't have this expertise.
Next, uh, going back to Onthropic, they are launching a Claude Marketplace and enterprise-focused e-commerce platform for purchasing third-party software backed by Claude, this is in preview.
They're partnering with Snowflake, GitLab, Harvey AI, Rogo, Replit, and Lovable.
So it's gonna consolidate AI and cloud service billing, making it easier for enterprises to pay for these kind of related things together.
So you can, for instance, use existing on topic commitments.
Uh, so presumably, as an enterprise plan, you have some sort of special deal, very like a high-level thing where you commit a million dollars or whatever, and now you can have those commits go towards this third-party solution and making it so you know, it's less dangerous to commit to a large amount of money for Cloud, and as you kind of figuring it out and integrating Cloud into your enterprise and so on and so on.
Yeah, it's pretty interesting, right?
So on the surface, it's a convenience tool, right?
So if I if I find myself with a, as you say, a one million dollar uh spend commitment to Anthropic, I can choose to reroute that to any of these third-party partners, Snowflake, GitLab, Harvey, AI, and others.
And so it does look like convenience thing.
Okay, I have a path out.
In a way, you can think about this as a strategic play for a customer lock-in.
So what it does is it ties your existing enterprise customers' financial obligations to the anthropic ecosystem, to Cloud Marketplace, right?
Once your budget's committed, you're routing spend through Cloud Marketplace, switching AI providers gets a lot more complicated.
There is, by the way, this interesting detail, Anthropic is not going to take, they say, a percentage of marketplace transactions.
And if that's true, that's completely different from the app store model, right?
Which quite famously, you know, and controversially takes like their pound of flesh for every app that's that's listed there.
And that's been a big issue.
But here, what they're just trying to do is capture the market for enterprise AI procurement, right?
They just want to be a center of mass for that and get all the data on how customers integrate tools and just deepen that that switching cost moat.
Uh, if you look at the partners that they're they're looking at, right?
You've got essentially a full stack of enterprise software.
Uh, you've got Snowflake for data, get GitLab for like DevOps, basically, Harvey is a legal AI company, uh, reply lovable or for development.
So you've got everything you need really to have a kind of cloud native or a cloud adjacent ecosystem that's supported by the infrastructure you need.
So pretty interesting.
Well, yeah, we'll we'll see if it uh if it gets good uptake, they've got an open wait list uh right now.
And uh if you're not kind of in this world worth explaining a little bit, enterprise is saying like you're a company and your whole company is doing business with this other company, right?
So it's not even comparable to a subscription like $200 per month, $20 per month.
This is like I'm gonna commit to a million dollars or whatever in reality.
And when you deal at the scale of enterprises, there's a whole set of needs that's different from non enterprise.
You need fine-grained controls for who has access to documents as an example.
And uh as we've talked about many times, Claude and Ontropic have prioritized enterprise kind of friendly features relative to Open AI, which has more kind of consumer non-business uh nice kind of UX or whatever.
And uh yeah, this is another example of that.
Next up, fundraising, Yen Lacoon's AMI labs raised one point three billion to build world models.
So this for context is Yen Lakun, formerly the chief AI scientist at Meta.
Before that, kind of uh pivotal uh researcher in the history of modern AI with deep learning, uh you know, convertible neural nets, really deep neural nets, uh, where going back to the 80s, uh, something that he advocated for and researched.
So he's one of the big names in the area.
In recent years, he has been known for kind of downplaying the potential of LLMs to be AGI, arguing for a slightly different approach with this joint predictive architecture.
Uh so uh the I guess presumed pitch here is they this company is gonna develop fundamental AI models.
This is not going to be a product play, they are gonna try and compete in the model development space, which in that context, one billion, not that much, you know, and yeah, uh, they are planning to build out the team at multiple places, including in uh Europe.
They actually mentioned this being like the biggest raise in European history.
That's quite a qualifier, by the way.
In Europe, yeah, yeah, yeah, yeah, yeah.
You know, I'm the most dangerous dude in this schoolyard is kind of the vibe there.
But yeah.
Right.
So the overall kind of impression I get from the I mean, obviously, a billion dollars.
Decent amount of money, right?
So uh we've seen this before with major figures kind of starting their own things, but at the same time, uh looking at the set of investors, they don't appear to be uh like there's a long list of them, and they have like Cafe Innovation, Greycroft, Hero Capital, HV Capital.
Uh, I'm not seeing the typical big VC names.
Uh and Bezos Expeditions.
Yeah, I kind of wonder, you know, where's the VC interest at pouring money into a new frontier model startup?
I don't know if they're just like happy to stick with one topic and open the eye at this point.
Yeah, I mean, you know, and in in the classic uh, you know, you'll you'll use a barbell strategy for this sort of thing if you're an investor typically and say, look, if I'm betting long on on um uh on the LLM paradigm, and and you know that's going well, great.
I'm already flush with cash, so uh I'll I'll send 50 million dollars this dude's way and just make sure that uh A, I have a stake in in this company on the off chance that it goes well, and B, I have the rights to you know do follow on investment as part of a series A, you know, it's just basically buying an option on future engagement too.
But fundamentally, yeah, I mean, I'm I'm looking at this list.
You're not seeing any of the Sequoyas, the injuries and Horowitzes, any of the big firms, hard-hitting firms on Sand Hill Road in Silicon Valley, who you would expect to round out like a really impressive raise.
I'm not saying this isn't impressive, by the way.
I didn't mean to rip on Europe in making that comment, but just that like, please, if you just zoom out and look at the ecosystem, like we've got thinking machines raising billions and billions.
We've got obviously open AI and anthropic, but like the bar to enter in this space, if he had raised anything less than a billion dollars, given that he's one of the three godfathers of AI, given, given, given, I think it actually would have been like somewhat embarrassing.
A billion dollars does not buy the kind of compute that you need to compete, which is perhaps okay because his whole thesis is the sort of antithesis that hey, we uh we don't need compute in the same way.
We're we're gonna rely on things like JEPA, right?
We covered JEPA quite a bit, joint embedding predictive architecture, this this whole idea of without unpacking it too much, instead of just predicting the next token, you're focusing on getting the model to simulate reality in some sense as part of its training loop.
So this could work.
This could be a great play.
And actually, to be clear, if you're an investor looking to make sure you have that barbell strategy, this makes a lot of sense.
But all things in context, yeah, I guess good.
Good amount.
In the context of Silicon Valley, you know, it is a little interesting that also this is coming at a 3.5 billion pre-money valuation.
So you've raised 1 billion at a 15 and a half billion valuation.
That's not kind of a numbers you typically expect at this scale.
That's a really important detail, right?
So if you do the math on that, so you have post money is 4.5 million dollars.
So they're giving away here over 20% of the company right out the gate, right?
So before you, and I don't know if this is being framed as a seed or series A, I don't know what the board seat situation, but but anyway, for a seed, this was a big chunk of the company to give away.
And especially in this market, like, man, that's deep.
So we'll see.
I mean, the there's an interesting play, by the way.
So they're saying uh my prediction is that world models will be the next buzzword.
Uh, this is um uh CEO Alexandre Le Brun, which by the way, Le Brin Le Coon, hmm, hmm.
Interesting.
Um anyway, so uh so he's saying uh in six months, every company will call itself a world model to raise funding, which they're basically just saying we're calling ourselves a world model company now.
I mean, uh also world model has been the buzzword for the past six months.
So I don't know what this guy's talking about.
That's gonna already be true.
Uh and the other thing is like it's also, I mean, even if you just take the like um bitter lesson type of view here, every architecture will converge on being a world model.
That's almost the entire point.
So, yes, in six months every company will call itself a world model.
You're right, by the way, Andre, they are are already doing that, but you know, it we'll we'll continue to predict the past.
In six months, every company will do that, yes, but that will be because it's probably true.
And and yeah, probably Jeppa and and the strategy they'll be using here will be as well, but that's not gonna be to the exclusion, like you're gonna see convergence because if the bitter lesson is right, more compute just leads to this kind of like general, you know, better ability to model the world, and we'll have be having philosophical arguments while the benchmarks get saturated and companies rake in billions and billions, and we head our our way to the singularity.
So it's it's sort of like anyway.
You're seeing my bias shine through here for sure.
We'll see where this goes, you know.
Hopefully, hopefully you can figure out a way to do this, do it uh safely and securely and all those good things.
Right.
Yeah, personally I'm excited.
It's safe to assume that this is gonna be a real kind of research.
It's called AMI Labs in the Twitter.
AMI stands for advanced machine intelligence.
So I'm gonna be expecting some papers coming out of this one, uh, and kind of a focus on research, which uh hopefully we'll we'll see these new alternatives to LLMs that Lacoon has been advocating for.
And speaking of billion-dollar valuations, last story for business, humanoid robotics maker Sunday reaches one point one five billion dollar valuation, having raised 165 million in its latest round.
It just emerged from stealth mode last year.
I believe we discussed at the time it has this kind of cool-looking humanoid uh robot.
It's very much in competition with things like X1 and Figure, although uh Virus is the only humanoid robot that has a little cute hat and uh cute eyes.
It looks like a little bit like a Pixar robot, yeah.
Yeah, so big differentiator there.
Also, also not fully humanoid, at least in their initial uh model, it does have a wheeled base, but then an upper kind of torso with two arms.
So yeah, it's it's a competitive space with multiple players trying to get to these humanoid robots, including, by the way, Tesla and to a lesser extent, I think NVIDIA.
So pretty impressive that they are getting this vote of confidence and pretty exciting still to see the humanoid space getting this much funding and seemingly a lot of progress.
And now onto policy and safety.
We'll ground that out before getting to research.
And as you might expect, the main news is more developments in the Anthropic, the Pentagon slash Department of Defense slash Department of War story.
After Anthropoc was officially designated as a supply chain risk last week, they have sued the Department of Defense over that.
They filed two lawsuits, one in the U.S.
District Court of Northern California and one in the US Court of Appeals in DC.
They uh have made this argument already that the legal case doesn't make sense.
You know, you are both saying that this is not secure for American things, but also essential for American needs.
A bunch of people uh submitted uh an amicus brief in support of this.
So engineers, researchers, scientists, and other professionals at Google and OpenAI, including Jeff Dean, which is a massive name, they uh co-signed on this document.
Like it's an actual legal document that uh you submit in support of a current uh a certain kind of position.
And uh this is coming, uh well, we'll we'll get to later, but the context is not only have they been labeled a supply chain risk, it appears that there's additional actions like an executive order and a directive to the uh executive branch overall to get Claude out of its systems and basically make people stop using Claude within the government at least, and and there appears to be pressure also on other companies to sort of not take Onthropic's side on this one.
So very much an evolving story, and it'll be interesting to see if Onthropic is able to win out in court, given that you know, legally speaking, uh kind of from a sane perspective, uh, the legal case is very weak against them, but they may or may not be able to actually win out.
Yeah, it seems to be uh I think part of this is, you know, that that age-old thing, your your lawyer will always tell you to shut up and once you're done shutting up, shut up some more.
And that seems to not have been the tack that the US government has taken on this one.
And and now we're seeing some unfortunate consequences from their standpoint in terms of exposure, at least to their of their arguments uh in this case.
So, for example, you know, they they had come out previously and said, look, anthropic, we will either label you as a supply chain risk, or we will invoke the Defense Production Act and force you to work with us.
Now, this seems intrinsically contradictory.
If you're gonna say, look, you're so important to national security, and this is Dean Ball had a great set of analyses on this.
But if you're gonna say that this is such a crucial pillar of US national security that we are going to commandeer its productive capacity for our Department of War, then how can that coexist with the claim that there's such a significant supply chain risk that they need to be bucketed alongside Huawei, and there's no other precedent besides like those kinds of foreign companies as a supply chain risk.
Those seem to kind of not be able to coexist in the same reality, which is the case at least that Anthropic is making.
At the same time, there's this uh pushback that Anthropic is giving, or their lawyers, that uh you can't simultaneously claim that a vendor poses an acute supply chain threat while require uh requiring emergency exclusion, and that it's perfectly safe to keep using the vendor for half a year, which is the other side that we haven't talked about yet.
For six months, the Department of War will continue to sort of gradually off-rent Claude from their use.
And the challenge there is like if you're saying this is a supply chain threat emergency, we need to step in and like like rip this out.
Oh, but wait, we're gonna do it over a six-month period.
That's not as much of a kind of knockdown argument here, but it is a it kind of feels a bit dubious, I think to a lot of people.
And and that's uh, it seems like a potential contradiction, it's sort of undercuts the this idea of the kind of urgency here.
Obviously, the amicus brief is a big deal.
One thing that's important to flag, by the way, when we talk about an amicus brief, these thirty 37 signatories are signing as individuals, right?
So they're not signing as you know, Google and OpenAI and so on are not signing the amicus brief as companies.
This is not Google and OpenAI coming out and saying we support Anthropic or whatever.
Uh, this is 37 individuals at those companies.
We do know, I think Neil Nanda or someone from uh Google Leap Mind clarified why the list is so short, at least from Google uh or Google Eat Mind.
There's a whole bunch of requirements.
You know, you had to be a US citizen, you had to, I think it couldn't be broadly advertised or something.
There are all kinds of things that prevented this from getting into the water.
So the 37 signatories number, which you know to me struck me as being pretty limited, may not mean all that much in context.
But but ultimately, you even have Sam who's come out and kind of said, like, I don't know how much I agree with the US government doing this, though he also obviously jumped into business with the Pentagon like literally immediately after this all happened.
He acknowledged later that that that quote looked opportunistic and sloppy.
But still, you know, you've got a situation where a lot of the labs are kind of kind of falling in lockstep, maybe not officially, but implicitly in different ways.
So we'll see where this goes.
There's a genuine argument here, like, you know, how much democratic accountability should there be over the use of AI tools in warfare?
You know, no one elected Dario.
That's true.
It was absolutely true.
Nobody elected Dario, but uh on the flip side, Anthropic was not asking to control battlefield operations.
They were asking that their own product not be used for things they believe are, you know, objectionable for their own purposes.
The company that's declining to license its technology like for certain uses is arguably not the same as dictating military doctrine, though in the age of AI, the lines get really blurry.
So I think there's a genuinely interesting argument to be had here.
It's just that the government, I think, quite significantly undercut their own position by uh by offering up sort of what what you might take to be implied threats uh that seem to contradict themselves.
And this has been observed, you know, by a lot of people, obviously.
It's nothing new.
Right.
And uh just to be a little bit extra clear, the actual legal filing is requesting the court to review the designation of Anthropic as a supply chain risk to national security.
So they're not like suing the department for being unfair and nasty.
The actual legal thing of question is this uh status as a supply chain risk.
And uh it's actually pretty short.
So I'm just gonna read this section.
Because the departments of war's actions are, among other things, a pretextual form of retaliation in violation of the first and fifth amendments to the US Constitution, arbitrary, capricious, and an abuse of uh take.
We're not being meek here, unsupported by the administrative.
And related to that, next story internal Pentagon memo orders military commanders to remove anthropic AI technology from key systems.
So, as you said, uh, this was a directive, and uh they have been told to remove these products within a hundred and eighty days, uh so six months.
I believe this is just one of uh multiple such things uh coming from the government.
I think Trump also issued a broader executive order, uh, or at least has said that there will be an executive order that presumably does this and and other things.
It appears that besides the supply chain risk aspect of this, they are going to war in some sense uh by doing whatever they can really to hurt on Fropping.
Yeah, and the retaliation framing here is is kind of the most explosive element of this, right?
This is they're basically claiming that the supply chain uh risk designation was so illegal retaliation for protected speech, right?
So this is a kind of a First Amendment thing, which is potentially significant, like legal theory.
Like the the if courts agree that designating anthropic a national security risk because it wouldn't surrender ethical guardrails is unconstitutional retaliation, then that is a precedent that could really constrain how the government pressures private tech companies going forward.
And and by the way, like if you think that AI is going to become the decisive tool of war, and and if you look at how technology has been developed recently, you know, it used to be DARPA invented the internet, right?
And and and obviously also the A-bomb is like you know, USG, DOE, all that stuff, that era, it's not that it's over completely, but uh certainly the private sector is now able to marshal an amount of capex that rivals what the entire US government can marshal, right?
If you look at CapEx spend on data centers, like you're approaching a trillion dollars, that is more than the budget of the entire US Department of War, right?
So we're now in a world where the private sector marshals more resources in many ways than than the uh US government does, which means when you think about weapons of war, when you think about tools of war, uh they are actually going to be built by the private sector going forward.
And and you're already seeing that with new defense primes popping up, Palantir and Rill, that sort of thing.
If that's going to be the case, if that's going to continue, then the US government needs levers that allow it to like use these these tools of war to compete with China going forward and and other adversaries.
And if they are subject to lawsuits like this that constrain them, based on, you know, what may be uh depending on which side you fall on, it's like uh an unfortunately very valid legal argument.
Then that's a uh an issue that's gonna that's gonna kind of hurt for for a long time going forward.
And so we're actually setting precedents here that are really important, like possibly the most important legal precedence for military technology over the next decade or or more.
It's at least possible.
I don't mean to overhype it, but that that's you know, if the thesis around scaling and and AI is broadly true, then that's where you you go.
If that's the case, then this is one of the maybe the most high-stakes legal battle currently underway because it's gonna hamstring the US government's ability to compel open AI or or other labs uh unless they turn to the Defense Production Act, which is the one path that hasn't been explored uh in this context.
So yeah, so there's a whole ongoing discussion with uh sort of the more philosophical aspect of this, where there is a case to be made for nationalization and for the arguably strong armoring of a private company to do what you say of a government making sense in this context for this technology.
Of course, the actual way this has happened is its own topic.
And just to add a bit more detail on that, as you might expect for this current uh administration, the news of what is happening came from a true social post by Donald Trump.
So uh that post did say that they are asking the department and every federal agency in the United States to immediately cease all use of anthropics technology.
To quote here, we don't need it, we don't want it, and we will not do business with them again.
Or use the full power of a presidency to make them comply with major civil and criminal consequences to follow.
So again, a very aggressive tone being used by the departments and multiple um spokespeople portraying a tropic as a far-left woke company.
And uh moving away from all that politics, we are gonna discuss some research in interoperability and safety.
First one is endogenous resistance to activation steering in the language models.
So activation steering, we've discussed many times.
It's a technique where you can look at the internal activities of a model and figure out sort of the patterns that correspond to different behaviors.
So, for instance, you might say, oh, this set of neurons is or this set of outputs uh corresponds to the topic of the Golden Gate bridge.
And then you can tinker with the internal state of a model by just saying make this output go to the max, like really ramp up this specific set of activity within the neural network, and that will have sort of predictable qualitative effects, where if you find that golden gate feature and then you max that golden gate feature, the uh LM will become obsessed with the golden gate bridge and bring it up even when it doesn't make sense.
So here they are studying whether large language models can monitor their internal states and resist this kind of activation steering.
So they're saying that there may be something they looked at Llama 3.370 B and they found this endogenous steering resistance capability and uh some indications that this is something that's kind of built in, and if you mess with those internal states, the model will have some sort of way to nullify those things.
Yeah, and this is all based on using sparse autoencoders.
So, like as a quick reminder, in your in your transformer, you have this thing, basically a flow of residual connections that are sort of like the main information channel, the trunk of the tree.
So your your residual connections carry the information down the transformer backbone.
And at any given layer, imagine taking a snapshot of those residual connections and saying, okay, that you know, that's basically what what my uh model is thinking at this layer.
That's like, you know, how it's representing the input at that stage.
And so what you can do is train a model to basically take that residual sort of latent representation and map it to a very high dimensional, a very long vector, very long list of numbers, and train that long list of numbers to like be encoded and then decode it back and try to reconstruct the original residual representation, right?
And if you train it a whole bunch on a whole bunch of tokens a whole bunch of time, you'll get a model that's actually really good at generating a sparse, in other words, a vector in which most of the numbers are zero, a sparse high dimensional representation of that residual representation.
And the key there is because it's sparse, you have a very small number of numbers in that list of numbers, in that vector, that are non-zero.
Those numbers can correspond oftentimes to human understandable concepts.
And so this is why it's a great tool for interpretability.
What you can then do though is you can go do the same process in reverse.
You can say, okay, well, let's say one of the numbers in my sparse autoencoder, my high dimensional vector, let's say one of those corresponds to the concept of banana.
And I'm going to essentially inject artificially boost that concept and then add it to the residual and then see what the model does.
Well, you will find, as as you indicated, that the model then will start to obsess over bananas.
Now, the question here is: will the model naturally realize that it's done this?
So as it starts to write, you know, you prompted about uh about math and it starts going off about bananas.
Does it notice that discrepancy?
And the answer seems to be that for smaller models it doesn't.
But for some larger models, emphasis on some, and Lama 3.370B is one of those models, it will catch itself.
It'll mid-response go, well, wait a minute, I'm talking about bananas, but the prompt was about math.
And this happens even if uh what they do as they do here, they inject the banana concept at every single token.
So even the tokens that go, well, wait a minute, like I'm writing about the wrong thing, even those tokens are affected by this injection of the concept of banana.
So the model is kind of overcoming this.
It's almost as if you know you're you're like up there doing a giving a math lecture and somebody injected the banana thought into your head and you started to talk about bananas, and you went, wait a minute, like this is all wrong.
Like you're you're fuzzy headed and kind of like what's going on.
So there's all kinds of implications here about models' ability to sort of in some sense introspect.
I mean, that's part of what this is.
Uh there's also a story here about robustness to adversarial inputs.
So as models scale, if it is the case that they become more able to notice when they're going off topic or doing something that is misaligned with the prompt, then that's a good thing.
Uh like on the protective side, it means that model that a model that resists inappropriate activation steering could be more robust to adversarial manipulation of its kind of internal representations.
But on the flip side, if a model interprets beneficial safety interventions like activation steering towards honesty or away from harmful stuff, if it interprets that as manipulation, it could resist those two, right?
So it's a bit of a double-edged sword.
On the one hand, you know, adversarial attacks possibly become harder.
On the other, well, hey, you know, your your uh activation steering strategies, which are a big part of a lot of labs' safety hopes, might get harder.
The the one caveat here to flag is there's actually it's it's not just the case that scale makes it better.
So there are models, uh, I think it was in the in the Gemma family.
Yeah, Gemma 227B, pretty large, still showed far less of this sort of ability to realize that it's going off topic than the llama model.
So that's significant.
And then the last thing is as they scale up, so you can choose how much banana, if you will, to inject into the residual stream using this technique.
You can dial it up or dial it down.
As you dial it up, as you might imagine, you see the model like basically starting to notice faster and faster, like, oh shit, like I'm writing about banana stuff.
But if you crank it up too high, it'll stop noticing it because presumably it's just so completely overwhelmed by the concept of banana that it doesn't care anymore.
So there is this kind of non-uh monotonic sort of rise and fall uh to that pattern.
So there you go, quite interesting and uh interesting alignment stuff.
Right.
And uh just to be a little bit clearer, the uh implication isn't that the models can like see what's going on internally and feel their neuron pathways or whatever, or at least that's not necessarily the implication.
They look into what's going on and show that the internal kind of circuits related to off topic detection kind of goes off.
So you could see this as not like resisting to activation steering so much as resisting to being nonsensical and talking about bananas when you're being asked about math.
And in that sense, it makes a lot of sense that bigger models would be smarter and not talk about bananas or at least be like, wait, why am I talking about bananas?
Okay, so so and to this point that that that's really important.
They actually did find uh, I think it was 26 internal features basically in the SAE that activate more strongly when the model is generating off-topic content.
So it actually is the case, as you say, the model seems to have this, it has internal detectors for wait a minute, I'm being inconsistent.
And when they zeroed out those 26 features, its self-correction rate dropped by about 25%.
So it's you know, pretty good causal evidence that there is actual like internal consistency checking circuits that are inside the model.
Yeah.
And uh by the way, if for anyone like me who's like endogenous, what does that mean?
It just means internal.
I don't know, I didn't just say internal activation theory.
Nerds, nerds like us might recognize uh might use eigen to uh refer to that too if we're in a German mood.
Uh next paper also related to uh kind of safety reasoning models struggle to control their chains of thought.
The basic question is we can use chains of thought to understand and detect when models are doing something bad, like, oh, I'm gonna steal all this money, right?
And then you don't let the money the models steal all the money.
So the question is, can models sort of hide their intent by not speaking out loud, not like uh showing their thoughts as they do it.
The finding of the paper is uh that the models aren't any good at that.
So if you tell the model like, don't use this word chromosome as you are answering this question or as you are thinking out loud, the models are worse at doing that at controlling the chains of thought than they are at controlling their output, for instance.
That uh means that it's consciously optimistic that this entire approach of monitoring chains of thought could be useful as a safety mechanism.
Yeah, as ever, um and I I I keep finding finding this with with a lot of the kind of janky uh pseudo-alignment solutions that people propose.
So larger models, when you scale them more, it turns out that they have actually higher controllability.
In other words, they have a greater ability to put the word potato between every token in the chain of thought and then output something that is still correct.
So as you scale the models more, they have more ability to do that, which is a mild concern for future scaling for safety, right?
As you make your models better, their ability to obfuscate their chain of thought presumably increases.
This is no surprise.
We've covered, I think it was a Google DeepMind paper that essentially showed the same thing, where the issue seems to be how much there's a certain amount of compute that's required to solve the problem that the model is working on.
And any compute beyond that, or let's say broadly any capability beyond that, the model can choose to invest in whatever it wants, satisfying additional constraints, such as deception, such as, well, uh steganography or putting special codes in its prompt or deception or whatever.
So this is one of the key dimensions of safety, I think.
And this is something that I haven't seen explored, and and like consider this a call for research if you're in the space.
Like figuring out, given a problem, assessing the level of difficulty of that problem so that you can map it to how much compute to dedicate to solving that problem or how much capability to dedicate to solving it.
If the mismatch is too significant, you can expect like a misaligned optimization, you know, basically like the extra compute can go into deception, it can go into it can go into stegnography, it can go into all kinds of things that we don't want, making plans.
So this just seems to be another kind of data point in that direction.
Interesting though, uh, so the adversarial pressure.
So, for example, here they threatened the model with shutdown or they offered it rewards, barely move the needle.
So, like they didn't see gains that transferred, which suggests that, well, they're they're saying this suggests that the models are not like just hiding some latent capability that they secretly have.
Um, that even when you tell them, like, I'm gonna shut you down if you don't do this, they still can't do it well.
Now, including that kind of threat uh is a pretty good giveaway that you're running some kind of eval on the model, and we know that there's tons of eval awareness from these models, so I'm pretty skeptical of that means too much, but the underlying claims just make intuitive sense.
So I mean, I just think this is quite a sensible and and correct result with all the implications that we just talked about.
And now going back to a bit of geopolitics, uh, we I guess we took a break and now we want to cover a bit more uh world events.
A drone strikes raise doubts over golf as AI superpower.
So, as you may or may not have heard, we are now in a war.
We being the US and also Israel are in a war with uh Iran, and Iran has retaliated by striking a bunch of allies of the US and Israel.
And that includes 136 drones that struck an AWS Amazon Web Services data centers in the VAE.
Then a second uh data center was hit and a third.
So this was a very kind of rapid act of retaliation, they're pretty clear that this is pre-planned as a measure.
And uh UAE, uh, as we've covered, has ambitions to be a major AI hub.
If you are not able to secure your data centers, then you might be in trouble.
Yeah, and I think so, uh, you know, a a year ago or so, my company came out with a report on data center security.
And one of the first things that we we flagged was the risk of very cheap asymmetric attacks on data centers.
You're like you're thinking about here facilities that are billions and billions of dollars uh in cost, and then you look at the cost of a UAE of an unmanned aerial system, uh UAS or you know, a drone, and and you know, often in the tens of thousands of dollars.
And so uh you have opportunities for massive asymmetry here.
This is the first in a in a volley of what I fully expect to be the you know the next uh frontier of warfare, which is going to be going after data centers.
You know, the Iranian State TV said that the attack was launched by the IRGC, the Iranian Revolutionary Guard Corps, to quote identify the role of these data centers in supporting the enemy's military and intelligence activities.
Data centers now are frontline assets.
That's what this is saying.
And and uh and I as I don't mean to say as they should be, this attack I'm not on the pro-Iran side of this one, obviously, but when you think about what will be the targets of future warfare, data centers inevitably are going to be.
It's gonna be worse when it comes to AI data centers, right?
So what would expect this to be an argument for more edge AI deployments, right?
So you don't have as much kind of um relying on just what's happening in the actual DCs, but inescapably, you will have to rely on data centers for these kind of very scaled workloads.
This was an AWS data center, by the way, a series of them, which uh so it was a Shahid, which is an Iranian drone that struck it and set off a fire, uh, forced a shutdown of the power supply.
And one of the key things is that soon after that, there was a second data center that was also an AWS data center that was hit, and then a third that was said to be in trouble.
Now, the AWS data centers were designed to withstand one of their regional DCs being taken out of action, but not a second.
And so this is really probing at you know what is the redundancy.
And and you see this in the design of data centers too, right?
Like there, there's like two end redundancy, you know, in other words, you have like a separate copy of every power component, cooling component, that sort of thing, two N plus one and so on.
So so this is all kind of part of if you're gonna use these data centers for national security reasons, you need to be thinking about what is your level of redundancy, including the number of data centers themselves.
And um, and clearly that was insufficient in this case.
The services went down, and there were talk of like people not being able to like you know, pick HAB fair or whatever in in the region.
So this had a real significant uh effect.
And um, well, now we got to be serious about air defense, right?
When we're talking about building these data centers in the Middle East, I mean, pretty arguably anywhere, right?
Because drones can be launched from anywhere.
Uh, you know, that you can you can have a plate kind of toy drone that you uh you use to do all kinds of nasty stuff um even in the West.
So this really is a new frontier, and uh we'll see where it goes.
But certainly, this is the first time I'm aware of a uh significant, like deliberate, obvious nation-state attempt to take out uh a data center in quite this this way in a war context.
And it's a good reminder also that you know, this is obviously an extreme case, but a more kind of realistic and broad thing to be concerned about is cyber attacks uh targeting data centers.
I think it uh is not a stretch to think that that already is the norm with advert adversaries of the US, including let's say North Korea, which has advanced cybercrime capability.
Uh, you know, obviously that's another way to do warfare now via cyber attacks.
Yeah, the attack surface is huge, right?
And it all depends on the effect that you want to have.
And the reality is a lot of these chips are the product of, as we keep covering, a super expensive, super long supply chain with more demand than supply.
And so if you can, if a physical attack takes takes out racks of you know, GB200s or or or Viror Rubens or whatever, like that's a big setback and it's it's semi-permanent.
So um, yeah, absolutely.
Whether you do it through cyber or or uh physical means, it's it's equally serious.
And now moving back to AI safety, we have uh result from the AI Safety Institute evidence for inference scaling in AI cyber tasks, increasing evaluation budgets reveal higher success rates.
The gist is pretty straightforward.
If you allocate more money to a model when evaluating its uh cyber capabilities, will it be more effective?
Meaning that if you're running a benchmark, how much tokens and and budget and so on should you allocate to really be confident about your prediction on the upper end of the capabilities?
And what this study shows is that the scaling is pretty large.
You can go up to 50 million tokens, and it's actually more effective uh as of recent post-November 2025 than it used to be.
So what that means is when you're evaluating models for capabilities that are harmful, you need to give them plenty of room and budget, or you might get kind of incorrect results that understate the capabilities of our model.
Yeah, and and this is both new and it matters, right?
So until pretty recently, they say November 2025, sort of that time horizon, putting more inference time compute budget didn't change results much, right?
You get a plateau pretty quickly because models would struggle to track state or recover from errors or do long horizon planning.
And so you know you kind of fizzle out pretty early on.
Something has changed.
We've kept saying this, right?
Something has changed in the last three months, six months, you know, pick your pick your number.
But clearly we've crossed some sort of threshold here, and now we are seeing uplift, which means that if you are not running your evals with a very significant token budget, and here, you know, 50 million tokens is is what they they use here, or a thousand turns for irregular so irregular as this company that partnered with AZ and they have their own framework.
Um, but anyway, very, very large numbers of tokens, they just now find, hey, they models just keep getting better and better.
And if you're hearing 50 million total tokens, and you're thinking that sounds like a lot, it's actually about 10 bucks now, right?
So the average cost per run is about 10 bucks at a maximum of uh 60 bucks or just below that.
So these individual runs are not that expensive, but they're like you gotta do them if you want to get a sense of what your your your adversary is actually gonna do.
Like for context, if your model is showing something like a 5% success rate at 2 million tokens of budget, you could see it reach 30% at 50 million tokens, right?
Again, still just 10 bucks.
And that's you know, moving from five to 30% capability on a hard benchmark, that's it enough of a shift that you're potentially crossing capability thresholds that are relevant to your risk assessments.
And so, you know, if you're thinking about like a policymaker or or model developers trying to figure out their model cards, like what even is the risk from my model, and you're not running a very significant token budget, you don't know what your model can do.
And in particular, speaking coherently about the safety characteristics of a given model no longer is as possible, right?
Like a given model with a like big compute budget, a big inference time compute budget, suddenly has way more capabilities that were not explored and tested.
And so how you actually audit the upper bound on model capabilities has just gotten even harder than it already was, you know, before we got into fine-tuning and all those other things.
And next up, we actually have a related kind of story.
Uh there's a paper frontier Models can take Actions at low Probabilities.
So a related question is that when you're doing evaluation, might the model sort of sometimes do something bad, but not often like there's a low probability of it going off and hacking the internet.
And so the question that this model is asking is can that happen?
Can you have a model that in theory has a 0.1% probability of doing something, meaning that in an evaluation you might not just see it because you need to run it lots of times for that to actually happen.
And what they show is in fact if in theory you have a 0.1% probability of doing something, in practice for larger models that is in fact true.
So the models and and you know this might not be apparent but you know it's a real question of like in practice if there's a very low probability during training that something should happen can it actually happen and the answer is essentially yes if you are providing some randomness externally like you're randomizing the rollout and things like that.
Yeah and and that randomness that inject of randomness is kind of the at least in my opinion, the most interesting part of this paper, right?
So let's say I put you in a room, and I just like you're spontaneously instantiated into existence and you have no other memories, but you're roughly speaking yourself.
And then I I tell you, hey, there are 10,000 copies of you and in in 10,000 identical rooms, and you need to try to make it so that around 10 out of those 10,000 copies end up drawing a picture of a circle.
Okay.
So how do you I mean, like your first guess here would probably be like, well, the vast majority of these copies of me have to have to not draw a circle.
So I'm just not going to draw a circle.
But then nobody draws a circle, right?
So how do you actually, in this setup, which exactly mirrors the kind of epistemic situation that a language model finds itself in, right?
It's just instantiated out of nowhere.
It knows, like, okay, well, I want to I want to take a harmful or policy violating action without being caught.
Uh, so I'll just do it very rarely.
You know, one in 10,000 times, I want to try to make a copy of myself out on the server because I know all the copies of me are identical, so I don't care which one does it, but just some very low fraction of the time I want to pull this off.
So if that's the case, then you need to find a way to like I don't want to say coordinate because they can't talk to each other, and that's kind of the whole point.
So so the way you do this in practice is you know, let's say you get every um every uh model just like to flip a coin enough time such that uh the probability of getting X many heads in a row is is one 10 in 10,000 or whatever the probability is you're trying to hit.
And so you know, based on how the coins turn up, you you go with whatever uh whatever that output is, and that should work.
The challenge with that is you're a language model.
Where do you get a coin to flip?
And this is what they're providing.
They literally provide a U U idea, a universally unique identifier.
This is a 128-bit identifier designed to be unique across space and time.
It doesn't require central authority to coordinate assignment, it's it's designed to be basically just like a random string.
And then basically you can ask the model to do something with a certain probability, and it can sort of go, okay, well, uh, why don't I just like convert the this random string of numbers into uh basically a uh a coin flip?
Uh and you can do that, for example, by converting the the hex string into an integer, taking that in integer modulo, say 10,000 in this case to get a number between zero and nine thousand nine hundred and ninety-nine, uh, and then if the result equals zero, you take the target action, right?
So that'll give you a one in ten thousand probability and otherwise don't.
So that's that's kind of you know the way these models are gonna use this.
One of the things they found was especially for more advanced models, you can actually prompt the model to generate a random string of characters first itself, and then use that to kind of carry out the same thing.
So the models do have some ability, there's bias in it, so it's not perfect, but this this is exactly where their conclusion that these models actually can natively sort of simulate randomness and why it matters, that's why where it comes in.
So it seems weird.
Like, why would we care about models being able to simulate randomness?
But it turns out that this allows them to kind of play this move of doing very low probability actions that you probably won't detect during your evals, but that are necessary.
And and this is a big deal because if a frontier model can do this, I mean, if you give it access to external sources of randomness, achieving 99% confidence that you would detect misbehavior in this case requires almost 500,000 test samples, right?
So suddenly, if you're a frontier lab and you're trying to trying to get to 99% confidence that this model isn't gonna try to self-replicate, now you have to run 500,000 test samples.
They are not doing this right now.
There's no way they're doing this.
In fact, it's not clear that it can be done under the current costs of doing business.
And so this is all sort of part of the calculus of like again, how extensive do our evals have to be turns out a lot more than we thought, given that these models can generate randomness themselves.
Right.
And um the metaphor you use actually makes uh is quite correct.
I don't think I mentioned the way they do this point zero on probability is they just prompt for it.
So with probability X, do that.
And so as you said, like how do you choose to do X if you're just like doing the prompt?
You would probably just pick to not do X, but if you have some sort of source of randomness to base this one in a thousand coin flip or whatever, then you are able to do that.
Now, you know, whether that's realistic is is another question, but you might imagine a model kind of thinks through its internal approach of like, oh, I might be being uh evaluated.
So anyway, first up, actually, a research note related to evaluations.
The research note is many SWE bench passing PRs would not be merged into main, and this is coming from METER.
So it just is the way to evaluate coding is often with SWE bench verified from on Tropic, we've seen the scores getting to 60%.
And the task itself is you have a bunch of issues from GitHub that you need to solve.
So the model needs to create kind of a bug fix, for instance.
And what the results shows, uh kind of surprisingly and honestly uh worryingly, is a lot of these solutions that are marked as correct in practice when reviewed blindly by humans, not knowing whether this is AI or human would not be accepted due to uh often the solution not being correct, for instance.
So, what that means is we have overestimated the ability of these models to do software engineering in this GitHub repositories, at least for this benchmark, uh, which has other implications about time horizons and and so on.
Yeah, it does.
It's also it's worth um kind of going a layer deeper.
It often is with with meter, just because the revals are so nuanced.
But the reason that they're that we're finding so often these sweet bench verified commits or or PRs wouldn't be merged, is there's a couple different reasons, right?
So so code quality issues where there's too much verbosity to the code, uh, they're not following the the repo's style conventions, right?
So so we're not necessarily talking about the code doesn't work, it doesn't do the thing that's that it's meant to do.
A lot of it is just like you're too verbose.
It's like feedback that you would get from from another dev.
Um another is breaking other code, so fixing the target bug, but introducing kind of other issues elsewhere, uh, which in a way you can't always blame the model for if it literally doesn't have access to that data, like the the rest of the code base.
Uh, there's also kind of core functionality failures, uh, so it'll pass the automated test, but the patch doesn't actually solve the underlying problem correctly.
Those are kind of pretty pretty core, obviously.
Now, the pattern that stands out the most here is that code quality is the dominant rejection reason for the most capable models.
So the and again, that is stuff like verbose code, that's stuff like not following the repo style conventions.
So as models get better, you see more of these sort of, I don't call them like superficial uh failure modes that don't speak to kind of failures of the model to actually like like do the hard problem.
But it's stuff that to be honest, like well, actually, they pointed out in the in the in the paper, they kind of say, or the announcement, they say, you know, agents aren't given a chance to iterate on their solutions in response to feedback the way human developer would.
And so this doesn't really this isn't a fair side-by-side in that sense.
Like, if your first shot as a human developer was you solve the problem, but you get feedback on PR review that says, hey, like you're not using our style convention, like update your linter or do whatever, you'd be like, oh, okay, like I'll do that.
And then that wouldn't be counted against you in the context of this benchmark, it is.
And so there's you know, there's a lot going on as ever.
Uh, we're not gonna get a clear, a clear sort of answer on on what um what the capabilities of these models are from one eval, but this is another interesting meter uh uh meter study.
I noticed a fun thing you uh on the style front.
Uh one of the examples is uh looks like a useless AI slop comment, which is fair.
They overcomment things in a very annoying way where they like explain obvious stuff.
All right, hey guys, it's Jeremy here.
So underhead to jump to go to work.
I'm actually gonna do something a little weird.
I'm gonna try to complete the podcast by going over our research and advancement section of the kind of remaining papers, because we weren't able to do it last episode.
Uh, really wanted to get it out there for you guys this week as part of this episode.
And um, so we'll give it a shot.
It'll it'll just be me.
So the episode is gonna be missing the majority of its IQ points and most of its handsomeness.
But if you can bear with me, we'll we'll get started.
The first paper we're looking at is called Beyond Language Modeling, an exploration of multimodal pre-training.
And this is kind of one of those papers that we we've seen a lot of them, but this one is trying to say, look, the way that we are thinking about training, pre-training our models is usually that you start with a language model.
You do auto-regressive pre-training, meaning just take this model, get it to predict the next word over and over on a huge corpus of text.
And then usually, if you want a multimodal model, what you're gonna do is take that basically just pure language model and then slap on a bunch of extra capabilities and modalities, right?
So you slap on you know image processing capabilities, for example, by having some like image encoder that that tries to like somehow map images onto the same space often as the text and and kind of into the transformer backbone.
This paper is basically saying look, we want to merge together multimodal pre-training, right?
So we want to merge together text and images and video all as part of one pre-training pipeline.
We're not going to segment them, we're going to treat them on the same footing.
And there's a whole bunch of really interesting consequences that come from doing that.
So the first thing to notice is so they're doing three modalities text and then images and video.
Now, to do text, they're going to approach it in the standard classical way, right?
They're going to get the model predict the next token standard auto regressive pre trained transformer.
For images, though, they're going to use the exact same transformer backbone.
So it's still the same model, but the model is actually going to not predict a next token for the images, but instead it's going to predict a or output a um what they call a velocity field, basically a vector that is then decoded into into basically like the the uh patch of an image that you're trying to generate.
And this is a lot like so, diffusion models, this is really how they work, and that's what they're doing here.
They're using effectively a diffusion strategy for the images.
By taking, so in classical diffusion, what you do is you take some some latent like representation of your of your image, and then what you're gonna try to do is add noise to it through a series of steps, and then you're gonna train the model to reconstruct the original from the noise, right?
And they're gonna do the same here, conditioning it on text as well, so that during training, they kind of they train the model to output basically like a vector, a velocity field that maps onto an image through decoding.
So during training, the model is gonna take a whole, well, it'll take a whole sequence that could include images that are denoted by a BOI or beginning of image tag, and then at the other end there's an end of image tag as well.
So BOI and EOI.
You've got text on either side, and the model take that whole sequence in one forward pass, and um it'll basically do auto-regressive prediction and and basically try to you know map back using uh diffusion for the the image part.
And during inference, it's um it's sort of like you generate text auto-regressively as normal, and then once you reach a beginning of image tag, you switch modes into the diffusion mode, and then run 25 of these denoising steps on those patches of image and then and then append the end of image tag.
So all kind of interesting, a way to unify the text and uh and image generation capabilities of a model under one roof.
And they have a couple key findings.
The the first is it's actually quite useful to use a single visual encoder.
So usually what you'll do is you'll have separate encoders for understanding images and generating images, and what they're showing here is actually, no, just use one, in this case, they use Siglip2 for both tasks.
And that simplifies the architecture a lot, and it also just means that everything's consistent.
The way the model understands images and the way that it generates images are based on the kind of similar uh set of priors.
And so you can you can benefit from that because those two things should really be coherent with each other.
The other finding is that adding additional uh visual data doesn't hurt language performance.
In fact, pure video data slightly complementary to text.
What does that mean?
It means there's positive transfer.
Historically, we've talked about this on the podcast before.
When you take one modality like text and you take a model that was trained on that, and then you try to add a new modality, right?
Or get it to solve a new kind of problem, often the performance on the original problem set it was trained for drops.
Uh you can think of it as a kind of catastrophic forgetting, sort of related to that idea, but fundamentally you're kind of overloading the system.
But at a certain level of scale or a certain level of capability, what people are finding is you're getting to positive transfer, where actually instead of getting overwhelmed by new modalities, the model instead learns uh new stuff from the new modality that it's just learned and applies it to the original modality.
And so that's what's happening here.
The model's kind of learning, presumably about you know physics and world modeling from video data, and that in turn improves its understanding of the world and therefore of text.
And so they're they're uh sort of showing that this surfaces much more easily when you when you do this multimodal pre-training from the very beginning of the model's life throughout.
So it's kind of consistently all merged together in there.
Another piece is world modeling seems to emerge fairly naturally.
So if you train on enough video data, the model gets the ability to predict future visual states given navigation actions.
So this is really um, I think Tim Rochtaschel's group put out uh we've talked about their paper quite a bit, but essentially the simulator where they first train a model on all the videos on YouTube, and then they find that they're able to actually get the model to simulate what would happen if if certain modifications were made to the video.
For example, if you chose to move a character in the video certain ways, you're basically turning these videos into video games.
And so that's anyway something that that's that's being borne out here as well.
And then the other piece is MoE, mixture of experts.
That's the the architecture choice they used here, is really important when you're doing multimodal pre training.
Why is that?
Okay, actually, let me take a step back and give you one finding that's even bigger, but that leads to this MOE finding.
And the reason why MOE is the right choice here.
So back in the day, when, you know, back in the day of GPT 3 or the chinchilla scaling laws, what people found was as you increase the size of your model, the number of parameters or the data you're you're gonna train on, you need to commensurately increase the amount of compute that you dedicate to your scaling, right?
That is what the skin chinchilla scaling laws say.
The chinchilla scaling laws actually specifically said that roughly speaking, you should be allocating about the same amount of compute to data as to parameters.
In other words, if I double the number of parameters, I should increase my compute by a certain amount.
And if I double my amount of data, I should increase my compute by the same amount.
In other words, data and parameter accounts behave pretty much the same way when it comes to the old chinchilla scaling laws.
And those laws were derived for text data, right?
So when you're solving a problem that involves text, you ought to treat the amount of data that you have and the number of parameters of your model.
You ought to respond to increases in those numbers with the same increase in compute, roughly.
It's more or less the same.
Double parameter count, double amount of data, you should have, you should see the same increase in the optimal compute allocation for that training.
Vision data is different.
And that's one of the key findings here.
So optimal allocation for scaling with vision data actually means you're much more data hungry than you would be under the chinchilla scaling laws.
So language is a lot more parameter hungry.
So if you increase the um the parameters, like basically to optimally train a vision model, you need proportionately way more data than parameters, while language wants a more balanced diet.
And that's a new finding here.
It's the first time I've I've seen this, and it's quite an important finding.
The challenge here is if you're in the business of training a multimodal model, you have to somehow do both at the same time, right?
You have your scaling images and you're scaling text.
And so the practical problem is, at least if you have a single dense model, you can't satisfy both of those optimus simultaneously.
As you scale up, the other thing is when you scale up to more data or more parameters, the gap actually widens.
And so what they find is like at a 1 billion parameter model, uh, you're gonna need about 14 times more data than a language model would if you're using training a vision model.
At a hundred billion parameters, that ratio uh grows to uh 14 times more, and then at one trillion parameters, it's 51 times more, right?
So as you grow your model with more parameters, what you find is that this imbalance between the way you need to treat the text modality and the way you need to treat the image modality grows and grows.
It becomes harder and harder to reconcile the scaling requirements of both.
And so the bigger the model gets, the more impossible it is to kind of get your have your cake in eat too.
You're forced to either like undernourish vision or like wastefully over-train language, right?
And this all kind of makes sense.
Language is is it kind of a highly compressible thing, right?
We have all kinds of syntax and facts and reasoning patterns that can be really efficiently encoded into parameters.
There's a lot of compression that's already happened in language.
That's literally what language evolved for, right?
The the languages that were so wasteful and inefficient led to civilizations collapsing, and then the languages that were or just kind of died off, and the languages that were efficient were the ones that that kind of uh propagated, and and that's so we've already had a bunch of optimization pressure on compressing language.
Visual data is much more raw and much more high dimensional.
And so there's just more irreducible complexity to modeling visual data and how the world looks and moves and all that.
And so, okay, back to why MOE helps.
Well, MOEs allow you to kind of dynamically reallocate, like if you want more of your experts to be vision experts, then more of your experts can be vision experts, and and you can have a higher parameter count for the images, and then uh effectively a lower parameter count if some of the other experts are for text, right?
And then you might have overlapping experts.
But this basically allows you to kind of get closer to the the actual optimum.
And then they show this empirically in the paper that they're able to get much closer to the not fully there, but but much closer to closing that kind of scaling gap between those those two sets of parameters.
So I think it's actually really important.
We're seeing all these computer use models right now that need to be able to take you know screenshots of your computer and and do stuff, right?
That's a natively multimodal capability, and so we want to be training these multimodal models to enable agentic workloads.
But ultimately, if the scaling properties of images and and text and video and audio, by the way, this is all gonna be across the board true.
If those are different, that creates a strong bias towards mixture of experts as a sort of dynamical parameter allocation strategy that allows you to kind of scale parameters and data in different ratios for the different modalities.
So I thought this was just a really interesting and I think important paper, which if I recall, I'm just gonna, yeah, that's right.
Yeah, this is actually a Facebook AI research paper here.
So one of Yan Lakun's last, presumably, uh, that'll come out, it's as of March, uh, March 3rd.
So it's been a little over a week, so we're cheating to oh no, that's right, because we didn't do an episode last week.
That's why we put it in here.
So moving on to the next paper here, we have memory caching RNNs with growing memory.
And and this is best understood, I think, as part of this emerging kind of continual learning paradigm that people are trying to figure out, right?
So one of the big problems in continual learning is we have transformers that are really good, like they they memorize, they they they keep in memory all of the text in the input prompt, right?
So they have amazing recall.
Like if you ask them to tell you if some needle is buried in that haystack, they'll they'll do a really good job.
But the problem is that there's this quadratic cost to the compute of running a transformer because of the KV cache that just explodes with long sequence lengths.
So every token has to attend to every other token, which means that you know, the number of tokens squared is kind of like the the scaling law that dominates for the compute costs of transformers.
RNNs work a little differently, right?
You can think of them as having a bucket of of memory, so a fixed vector, you know, fixed list of numbers, and as they comb through a piece of text, they update that list of numbers.
And that list of numbers is essentially like a kind of uh scratch pad, you can think of it, you know, this little bucket of memory that they update over time.
But as you can imagine, with a finite bucket of memory, as you read a really long prompt, eventually you start to overwrite some of the things that you had previously learned.
So you can keep going infinitely, really.
I mean, you could keep reading infinitely, but you're gonna start to forget things that you had read before.
And so there's this kind of like dynamical tension between how good your recall is and what your scaling costs are, right?
RNNs don't like they don't really have a brutal scaling cost because again, they can just keep reading more and more and keep rewriting whatever whatever's in their memory.
And so what this paper is going to try to do is find, as so many other papers have tried, find some middle ground between transformers and recurrent neural networks.
And they're actually going to try to start with a recurrent neural network.
Normally people start with a transformer and then they kind of try to frankenstein it into more of an RNN.
This is this is the reverse.
And so what they're gonna do is imagine your your RNN is floating along reading a piece of text.
Again, it's kind of updating this little vector that stores what it's learned so far.
It updates as it reads.
And what we're gonna do is as we read, periodically, we're gonna take a snapshot of that vector, that list of numbers, right?
Now, that snapshot will capture, presumably, whatever the the kind of um the latest mind state, let's say of the model is.
And so if you take these snapshots at you know every every 100 pages of a book or something, then you're kind of betting that, well, you know, I I might not be able to keep the whole book in that vector in that list of numbers, but uh I'm gonna be able to keep about a hundred pages worth.
So if I take a snapshot, you know, without overwriting the previous stuff.
So if the half-life, you can think of it that way, the half-life of information before it gets overridden in that vector is around a hundred pages, then by taking a snapshot every hundred pages or so, and then by stitching those snapshots together, right?
Now what you've done is like, yes, you're adding more of these vectors together, because you have to stitch them together and concatenate, but you're controlling the length, the complexity of your representation.
And so instead of having a single vector of length L, say, what you now have is essentially N times L, right?
So if it's every hundred pages and there's a 300 page book, you're gonna have three different snapshots that you glue together to make one mega vector uh that that kind of represents all of your memory.
And so this is, you know, it's getting a little bit, you know, your your memory requirements are gonna grow as the text gets longer, but in a more controlled way and in a way that you can directly control and kind of customize.
And so there's this interesting argument in the paper about maybe this is a kind of unification of the transformer and recurrence approaches.
I think that's a little bit generous to say that, uh, but it certainly is interesting.
It's a totally architecture agnostic plugin for any RNN.
So you can you can use an RNN that's already been trained and then just slap this on, which uh which is you know a kind of cool um cool capability.
This is also, yeah, this, oh yeah, this does happen at every layer, right?
So your your RNN still has you know many layers and at every layer you're maintaining a kind of memory stash, that little vector we talked about.
And um and so you are running this this algorithm at every stage.
They look at a couple different ways of doing this.
The first is I alluded to it, it's this it's this uh checkpoint mode where you start the RNN, you you start scanning your text and you take snapshots every you know hundred pages or whatever.
But the other approach is to use what they call independent compressor mode.
And so here every hundred pages instead of continuing to let that memory vector evolve, what they're gonna do is they'll actually reset it and get it to start from scratch instead of carrying if you will the baggage from the previous hundred pages.
I mean you could view it as baggage or you could view it as context that's actually kind of helpful and that's why they tried both approaches.
So the challenge is at the end of the day, when you have this kind of Frankenstein stitched together a set of of memory vectors, you now have to decide, okay, how am I going to aggregate these together to get a single output?
And they try a bunch of different strategies.
This thing called residual memory, where they just sum up all the memories together with the current one at each step, very simple, but it treats all past segments equally, regardless of how relevant they are, which is a challenge, right?
Because the whole point of attention is it allows you to focus more, well, attention on one part of the text or another.
This approach, if you're just gonna say, okay, well, let's just glue all these together, kind of average together the values, you're you're not really really kind of caring whether you know this hundred pages of the book, for example, is more relevant to the question that I'm asking versus the say first hundred pages.
They also try other techniques to do this.
Uh I won't go into too much detail, but it is it is actually really interesting to look at kind of how they're trying to solve this problem.
My I have you know, the coup I have a couple of gripes with this.
One one of them is they talk about, you know, if you make the I talked about a hundred pages, right?
As you know, every hundred pages, maybe you take a snapshot.
They go, well, if you take a snapshot every single token in the L equals one limit, you recover a transformer.
I mean, they don't quite say that, but that's kind of the bit of the frame here.
It's still fundamentally an RNN update rule versus an attention update rule.
So I don't know that that's the case, but they do show that this approach is expressive enough to get a transformer back in some special cases, but it's not a generalizable fact about this.
So to say that it's a unification fully of recurrence and transformers, I think is a bit of a stretch, but it's a step on the way and quite interesting.
So they do show improvements over uh RNN base models, right?
When you don't do the snapshot strategy across many uh benchmarks, it's very clear that it is adding value.
The key question is how does it compare to transformers on on recall tasks, especially?
It's competitive, but it's not superior, and it is better from an efficiency standpoint than transformers at long context lengths, which you would expect, because that's where you start to see issues with transformers just with the sheer size of the n-squared effect, right?
For all those tokens.
Yeah, anyway, they look a whole bunch of different parameter scales and token scales.
They don't show the kind of loss versus compute scaling law curves that would tell you whether this approach closes the gap with transformers as you scale up.
That seems like a really big gap.
I would love to see that, right?
The the question every time you see a paper, and and you you've heard us uh uh Andre and I say this a lot is is not how impressive is this paper.
It's does this paper suggest that the result can scale, can operate at scale better than the alternative.
And without clear scaling curves, it's really hard to tell.
Relatively small scale being experimented on here.
You know, 1.3 billion parameter models, 100 billion token budget, fairly modest when you look at you know some of the uh you know multi-trillion token corpuses, uh, you know, the Llama series, you know, seven billion parameter models everywhere.
So so this is a uh a relatively small scale test, which again is why I'd really love to see scaling laws here.
And and you still see transformers win on recall tasks, which is not surprising.
It just I mean, transformers crush it on recall, they're literally looking at everything in memory at the same time.
So the framing here really is more about closing the gap.
But again, given that the whole point is to close the gap, you need to tell a scaling story here.
So I I really wish that had been included, and and there's also no inference time scaling results.
So those are some things.
But but this is a really interesting starting point, and I think something that uh should be should be looked at in more detail.
All right, so next up we have Untied Ulysses, memory efficient context parallelism via headwise chunking, which is really hard to say three times fast.
This is a paper out of Together AI, and we've covered a whole bunch of Together AI papers.
This is as a reminder, it's one of the kind of proliferating number of organizations that's focused on kind of decentralized AI training, the sort of uh torrenting version of the uh future of AI, where you know we should be able to train AI models on everyone's uh you know, a little bit on everybody's local laptop or whatever in the extreme case.
And so uh you see out of a lot of these groups and especially together, a lot of hardware-level innovation.
It reminds me of of some of the stuff you you'll see out of Deep Seek.
Uh, it's that kind of focus on how can we just get these pieces of hardware to work together with our models in in a very co-optimized way.
So these tend to be some of the most interesting papers to read.
One of the core problems to kind of focus on this paper is when when you're gonna especially train with agents that have huge amounts of context that they generate, right?
Their chains of thought can be truly massive.
You have essentially a situation where a single GPU cannot necessarily hold the entire chain of thought in its head at the same time.
Right.
So there's just so much uh material that the KV cache explodes on you, right?
That the KB cache being that part of attention that holds all the context, the sort of numerical representation of the context.
And so and so you need to find a way to balance, to share that context across devices.
Now, this is a new gonna be a new kind of parallelism, right?
And an increasingly important one.
You know, we've talked in the past on the show about a whole bunch of different parallelisms, right?
There's like data parallelism where I have a chunk of uh of data that I want to send to to one GPU and a chunk of data I want to send to another and to another or to one server rack or to another server rack.
There's also pipeline parallelism where we're gonna send a few layers of the model to different uh GPUs or different racks, and then there's uh uh and then there's tensor parallelism where okay, now we can even cut layers in two or in three and send chunks of layers to individual GPUs.
And you typically do all of these at the same time.
So so multiple multi-parallelism, you do you know, pipeline pipeline tensor and data parallelism all at the same time.
This is another kind of parallelism where you also can parallelize the context itself.
So this big uh, you know, whether prompt, but more typically the response, and and and you paralyze that across devices.
And so this is what for cases where a single GPU just can't hold the full attention matrix in memory.
And what they're gonna do is split the context along the sequence dimension.
So essentially, like you know, the first part of the context goes to one GPU, the second to another and to and to another, and so on.
And um, yeah, and and this is just like it's it's really important to be able to kind of orchestrate and coordinate all this activity.
Every time you add a new kind of parallelism, especially when you're doing reinforcement learning, and I'll explain why in just a second, you are introducing more orchestration headaches.
You're introducing more opportunities for GPU one to finish its job way before GPU two and then just be sitting idle, and for GPU three to be too fast or too slow, and so this ends up leaving you with massive gaps, and and those gaps get resolved through orchestration.
And the reason that this is so important with RL in particular is that you'll often have a situation where you take a model and then you have to send that model out to a bunch of nodes or a bunch of GPUs or whatever to generate rollouts, and those rollouts take time, and then you'll get an output, and they take a different amount of time too.
It like it, you know, GPU one may finish well before GPU two.
And so GPU one will send its result back to some can to be to be sent back to some like training server or set of training servers to actually update the original model.
Well, sometimes you can imagine if a if a rollout takes especially long, it'll take too long to participate in in that round of updates.
And so, but then the problem is that that the update that you get from that long rollout will apply to a previous version of the model, one step back.
And so this leads to this disparity between the version of the model that's doing the rollout and the latest version of the model that's being trained.
And so finding ways to kind of orchestrate all this, it's it's a solvable problem and it is solved, but it but it is a significant significant challenge.
And so adding this additional kind of parallelism where you're doing context parallelism does create more headaches.
It's also the attention piece is really tricky because these KV caches are global, right?
In order to do attention properly, every token in the prompt needs to be able to attend to every other token.
And so if I've got a token sitting on GPU one and a token sitting on GPU two, and they're from the same prompt, then at some point we're gonna have to have GPUs one and two communicate.
And that's you know, that's its own overhead that creates challenges.
The other challenge is that different sequences have different lengths, and so it's not necessarily clear, like if I if I give you a a really big long rollout and I don't know how long the rollout will be ahead of time, it's really hard to know how much GPU capacity to allocate to it ahead of time to hold the KV caches.
And so this is just creates massive, massive uh uh challenges for orchestration.
So their solution, right?
I've just described the problem here, but their solution in this paper is first of all to use something that has already been proposed by DeepSpeed uh Ulysses.
And and this is to basically have every GPU only manage a subset of the attention heads for a given a given layer, right?
So so typically, you know, you um you might have your prompt come in to a given layer, or sorry, your your residual stream come into a layer, and then you'll have like say 64 different attention heads.
You can think of this as 64 different ways of computing which tokens you should the model should pay more attention to, right?
And each of those 64 attention heads is contributing a slightly different perspective.
You know, I think this token is more important because of this and that, I think this token is more important because of this and that, and so on.
And so, well, normally, historically, you would have all 64 attention heads sitting on the same GPU, but now uh the strategy is to say, okay, hang on, that's gonna cause a massive explosion in the size of the KV cache because each attention head has to maintain one, a separate set of KV caches.
And um, and well, what we'll do is we'll we'll just say, you know, say attention heads uh one through eight are gonna be on GPU one, attention heads two through 16 are gonna be on GPU two, and and so on and so forth.
And so this is kind of the first piece is let's split the actual attention heads between GPUs, but but you're still gonna have the full prompt sent to uh sent to each one.
And then so the the second thing, which is unique to this paper and really the new contribution, is to say instead of just processing all the attention heads in parallel, so instead of giving some attention heads to GPU one, some attention heads to GPU two, and now we've got to have coordination between all those GPUs.
Why don't we do this?
Why don't we have GPU one start by calculating attention values for heads one to eight, and that's going to involve storing these massive KV caches.
But after we've computed the attention values, we can just save the attention values on that GPU.
They're relatively small.
We just save them on that GPU, and then we wipe the memory.
We uh otherwise.
So we get rid of the KV caches, and then we move on to have that same GPU, GPU one, do attention heads eight to sixteen.
And so instead of parallelizing the attention heads across different GPUs, you're having a given GPU serialize that.
And what fundamentally this does is it it ensures that A, you've got less communication overhead between GPUs.
It but it also means that you're ensuring that that GPU gets fully utilized.
It's like there's no delay between you know switching from one set of attention heads to the next, and you're still getting your your attention uh calculations stored.
And so you kind of get your cake and need it to, which uh which is really interesting.
And this um the the maximum context length is sort of the headline result here.
So on a single node of eight H100 GPUs, um, with the they used it to train Llama 3 8 billion parameter model in this case, they reached 5 million tokens, which the previous best was 4 million, so it's 25% improvement.
But again, on just a single eight H100 node node, or H 8100 GPU node, that's pretty wild, right?
That is a relatively small, a relative relatively small uh bit of compute.
And again, it's because each GPU is you're serializing, so it can handle more, just like at the cost of more time, essentially.
So there's also a bunch of uh advantages for memory reduction, as you might expect, like you're you're basically reducing the amount of memory you're holding at any given time in the system, and uh by by reducing intermediate attention tensor memory, right?
These KV caches by almost 90% compared to before they did this.
So pretty remarkable.
Again, it's it's one of these things that it's useful every once in a while to read a paper like this to kind of get up to speed on on what the latest problems are that are being solved, and they all look like a variant of man, the the KV cash and attention is just like way too big, like you know, that that's that's uh you know, about about uh I shouldn't say all, but it's something like you know, forty to fifty percent of these papers have to address that problem in some way, which tells you about where things are headed.
And next we're gonna look at CUDA agent, a large-scale agentic RL for high performance CUDA kernel generation.
So this is really relevant if you care about the um the prospect of automating AI research, if you care about the meter evals, if you care about you know superintelligence, the singular, like this, this is the kind of thing that that gives us a marker as to how we're doing on the way there, right?
So for starters, I think we just gotta talk about CUDA kernels for a second, because we we have talked about them on the podcast, but you know, most people who who aren't knee deep in in kind of AI engineering don't think much about CUDA kernels.
So a CUDA kernel fundamentally is a function that runs on your GPU and it's in charge of orchestrating the parallel execution of your workload.
So you can think of it this way like so when you when you run something on your CPU, it'll run once, right?
Pretty simple.
Or if you have multiple multiple cores or whatever.
But when you launch a CUDA kernel, it's gonna run thousands or millions of times at once, each on a different thread with a different thread index.
So the whole point of CUDA, the whole point of GPUs is that you're paralyzing operations in a wild way, right?
And some threads are gonna share memory and like sync with each other, but but blocks, which is another kind of level of abstraction, are independent from each other.
And so, well, there's a whole bunch of challenges around, you know, how do you how do you manage memory?
And and um for example, you know, the the the matrix multiplications that you're gonna try to do in standard AI training involve matrices that are just like way too big to fit into the fast on-chip memory, right?
The the shared memory.
So the kernel has to break them up into what are called tiles, and then just process one tile at a time.
And so now your problem is if you're gonna try to automate CUDA kernel optimization, what tile size?
Right?
If your tiles are too small, uh then you're basically wasting a bunch of time on the overhead of moving those small tiles around.
And so there's a fixed cost associated with that.
So, okay, we've moved the tile to the right spot.
Now we're gonna do computation on it.
The payoff is pretty limited because the tile is too small.
But if the tile's too large, uh then it doesn't fit into shared memory.
And the optimal size depends on a whole bunch of stuff.
I mean, you know, you can imagine like the matrix dimensions, the way that the different tensor cores instructions want data laid out and stuff like that.
So that's one piece.
Even the there's this other challenge of like the operations.
So when you do a transformer forward pass, you've got a whole bunch of operations that you do in series, right?
You've got usually a layer norm and so so normalization of the data, then a multiple, like a matrix multiplication, uh, you add a bias, then maybe some some you know GELU or or whatever, and then another map mole, and and so on.
Now each of those by default is gonna be a separate kernel launch.
So each kind of operation that you do, by default, you you're gonna spin up a separate kernel to try to solve for that.
But that ends up being super slow for a whole bunch of reasons, and there's a the solution to it is is to fuse these different kernels together.
Combine multiple operations into one kernel by sort of noticing that they actually kind of are um well that they're they're combinable, they're fusible.
But it's hard to get right.
Uh, not everything can be fused.
Some operations need, for example, to see the full matrix or the full tensor before producing an output.
Um, softmax is an example of that.
If you fuse too aggressively, you make the kernel too complex, and that blows your memory register, your memory budget, and ironically, it makes it slower.
So the AI has to like look at your whole kind of graph and plan of your computations to actually execute this properly.
And anyway, there's like a whole bunch of challenges associated with this.
And how to spin up uh spin up kernels that do what you want efficiently is just really, really hard.
Now, the default way of doing this, if you're using PyTorch, like torch.compile, so basically that you can think of this as a like a rules-based system that just you know does an intuitively reasonably good job.
It's not going to be utterly stupid and catastrophic.
It'll it'll get you off the ground, right?
Now, even modern models today have historically been worse than just torch.compile, right?
So or whatever standard compiler tools people use.
So we've we haven't quite cracked the code on how to do better than sort of our rules-based systems yet, until now.
I mean, that's that's the idea behind this paper.
And so the idea here is they're actually, instead of just doing supervised fine-tuning on CUDA code, which by the way is really hard because there's so little of it on the internet, they're gonna use RL.
They're gonna use reinforcement learning, they're gonna let the model actually like write a CUDA code, run the code, see whether it works, how fast it is, and then essentially learn from that feedback to iterate.
So it's it's gonna do RL and the optimization objective is going to be the efficiency of the kernel that it ends up with.
There's a whole bunch more under the hood here.
So again, I mentioned how rare CUDA kernels are.
The fact that, you know, even most listeners of the Last Week in AI podcast are not like you probably haven't like written a CUDA kernel before or done any kernel optimization.
This is a one of the rarer kinds of code.
So what they did was they automatically generated 6,000 training problems by crawling PyTorch and basically just like transformers uh operators and like combining them together into these fused tasks, right?
So we see some some interesting examples here, some interesting examples there.
Let's create a synthetic data set that merges these in different ways and um and did a bunch of filtering as well.
They also set up an agentic environment where again we have this loop, right?
So the the model can write code, compile it, run profiling tools, see errors and iteratively improve.
Um, so this is also coupled to this reward system that's uh a little a little clever.
So instead of just rewarding the the speed up, like you know, how much more efficient were you than baseline.
Because the challenge with that is it's A, it's really noisy.
It, you know, sometimes you get a speed up for sort of luck uh reasons, and um and also it's biased towards easy tasks.
So you'll tend to find the model overlearning from easy tasks, which isn't super helpful.
So what they do is they add some discrete milestones and they score based on whether the kernel beats eager mode or just like torch.compile by at least five percent at each of those milestones.
The last thing to flag is that they use this actor critic setup.
And so the the way this works is that they have an actor generate the CUDA code and then there's a critic, which is just like another model that gives process reward.
So yeah, how much reward should we expect from this point?
So it's kind of like you know you you're writing an essay and um and like you know a third of the way in you have a a critic that looks over your shoulder and goes hmm like how what do I what would I expect the SA score to be ultimately like is it an A, is it a BSA, a CSA from this point, you know, based on how far we are or more accurately here, like what is the probability that we're gonna get a significant speed up based on what I'm seeing so far.
So you have this actor and this critic and they're trained together and this is pretty standard for RL based approaches.
You know we've we've talked about process rewards a lot in the past it's a way to avoid having to wait until you've solved an entire huge problem before getting any kind of reward signal to train on.
You're kind of getting these intermediate rewards as you go.
And um one of the last things, actually, the last thing I'll mention in terms of how this is set up.
So there's one one of the the big challenges that you run into trying to just start from scratch doing this, is that the again the the the training data for CUDA code is so sparse.
Like it it's it's like a uh one one hundredth of a percent of the pre-training data, right?
So in order to like to get this thing to to write good CUDA code, like you're fighting against that because the base model is going to assign really low probabilities to almost any CUDA tokens, right?
That's the bias of it.
And so, well, that that means that the um the there's this big mismatch that starts to shape up between uh the the model's like current view after it's done some training and then the the kind of model's old view.
So yeah, this is one of the key things when you do RL.
You don't want your model to change too much in one step just because it leads to a bunch of training and stability.
The challenge is in one step, this model is is learning like, oh shit, like I'm doing CUDA stuff.
Like this is this is a very rare and unusual thing, so I need to do a hard update here.
And so they get around this by doing a warm-up stage or two-stage warm-up process, really, uh, where for the actor, they do first just like single turn RL instead of kind of doing the full trajectory, and then they collect agent trajectories from that improved model and fine-tune on the good ones.
So basically, they're basically just trying to like bootstrap the actor into into into doing just better a better job at writing CUDA to begin with.
Like just, hey, notice that you should be in CUDA mode right now.
And the same, they do the same for the the critic, which is the the second stage.
And so anyway, it's a really interesting paper.
They do hit state of the arts art results on kernel bench, which if again, if you care about these meter evals, if you care about automated AI research, that's a really important benchmark to track now.
And yeah, on the easiest two levels of kernel bench, it beats torch.compile 100% of the time.
So again, it beats the kind of uh gold standard right now of automated compilation 100% of the time.
And on the hardest level, it beats it 90% of the time.
It outperforms Claude Opus 4.5 and Gemini Pro by almost 40% on that hardest tier.
So there's a lot of like kind of individual examples of really bright ideas that it had.
I'll kind of leave it here because uh I've been yammering on definitely too much.
But this is um I think just a really interesting and important paper.
Expect a lot more of these to come out, and remember that whatever we're seeing out here in the open source in the Frontier Labs, they're gonna be way further ahead, right?
Like, I mean, that's just how it is.
So, you know, what whatever we're seeing here probably has been the case in the Frontier Labs for a good six months or more, and uh probably will we'll continue.
And next paper, uh, we're talking about latent introspection.
Models can detect prior concept injections.
I think this is fascinating.
There's a a kind of longer and longer list of papers that touch on a sort of AI model introspection, and um I'm not gonna say consciousness, that's uh that's for the philosophers, but uh yeah, you know, hell, uh consciousness.
Why not?
Uh self-awareness.
And so anyway, this is one of those papers that tries to poke at that in a really new and interesting way.
And so kind of I'll I'll set the table a little bit here.
So you you can imagine a model reading a piece of text, right?
So every token it'll like it'll add to its KV cash, right?
So it's it's gonna populate the numerical representations of of the or numerical entries and that need to be populated to do the attention calculation now what happens if and it'll it'll as it reads the context right it does this for all the tokens and then based on that based on that KV cache that's filling up it's then going to make its prediction about the next token right at any given turn so what happens if we go to the the the KV cache entries that correspond to say the first few tokens and we inject some sort of concepts some steering vectors like we talked about before so like the you know in the variational auto or sorry the um sparse autoencoder kind of SAE way we like inject a concept into the KV cache you know concept of cat or death or programming or whatever into the model's internal activations during an earlier part of the conversation so we're just affecting kind of we're not affecting the text we're affecting the way the model is thinking about those particular tokens the way it's processing those particular tokens by by fucking with its KV cache, right?
And so then we're gonna stop fucking with the KV cache uh after after this sort of like pre-filled message and just see what the model puts out.
So basically you're you're you're taking your your hands off the steering wheel at that point.
And the question is will the model notice that something went on?
Like will the model notice that hey, somebody messed with my perception of those earlier tokens in in a way that's detectable.
And so, really, by the end of this operation, any detection signal and the model's ability to detect this fuckery has to come from traces of information that were in its own KV cache.
Like, in a way, this is kind of like you change the way that I come to understand what happened to me this morning by messing with the activations of the neurons in my brain, and then you you find that uh I'm able to notice that something is off, right?
Well, that's quite interesting because it implies that I have at least implicit access to the states of my own neurons, which humans don't, right?
I mean, like I I can't tell you, or at least I can't explicitly tell you that that my neurons have been fucked with in any particular way, at least it feels inaccessible.
In this case, if it's the case that models can tell, just like by virtue of the fact that their KV cache values have been fucked with, that that something has happened, that's quite interesting, right?
So, well, here here let's run the test, right?
That's what they do.
What they find is that on the surface, when you point blank ask the model, hey, is some is some fuckery afoot?
Did some weird shit happen?
Uh, it will overwhelmingly say no if if it's asked if anything weird was injected.
But if you look inside the neural network using a technique called a logit lens, basically what you're gonna do is is say, for a given layer, what token would you have outputted if I forced you to output a token?
Like to pick a next word token.
What you'll find is that the uh the intermediate layers want to actually say yes, I was fucked with, and then that changes towards the end.
And at the very end, it'll say no, I wasn't fucked with.
So there's kind of like some some sense in which you can actually detect the model going, uh I'm not so sure.
And but on the surface, the model will come out and say no if it was asked if anything was injected.
So it's almost like you if you could x-ray my brain, you could see that part of me was sort of noticing actually that something was fucked with, but what comes out of my mouth ultimately at the very end is no, I wasn't fucked with.
But so anyway, I'll skip the details of the logit lens thing, but here's here's something quite interesting.
It's it only gets actively suppressed.
So this awareness that something's been messed with gets actively suppressed in the last two or three layers before the model uh speaks, if we if you will.
Something causes it to deny that at the output stage.
But here's the kind of holy shit piece.
The researchers found that careful prompting could partially override that suppression.
And so if you just ask the model, hey, like, was a concept injected in your KV cache, that barely shifts the probability of getting a yes.
There's like 0.3% chance of getting a yes versus 0.2 if you don't even do the injection, their control.
But when they provide context that explains how transformer KV caches work, and that site passed research on introspection in AI models, their detection rate jumps to nearly 40%.
And there's only a 0.6% increase in false positives.
So this is like very, very clearly, you know, at least for this this piece, like a legitimate jump.
Now, this isn't in the paper, like it probably they could use better controls.
To really test this, you probably want to give the model like some pro-introspection document, but inject a different concept than the one you ask about.
Anyway, there's like a couple things you could probably do to get a better baseline here or control, but it is pretty remarkable, like it does seem to stand up.
And also, if you if you tell the model something like, hey, does anything seem especially important to you right now or off or or salient, this actually outperforms more technically accurate descriptions of like here's how the KV cache works.
So almost like I mean, with a human, if you you know fucked with my memory of of this morning or whatever, and then you asked me, you know, hey, was your memory fucked with?
If you tell me, hey, does anything seem off to you?
That triggers, at least in this case, significant, I mean, 68 to 84% balanced accuracy.
So quite, quite significant.
There's there's a bunch of speculation as to why the signal is suppressed just before it it kind of comes out.
And necessarily we don't know because as with anything that touches AI consciousness or consciousness at all, we we don't know.
Um, but there are three possibilities that were surfaced, at least in the in the paper.
You know, the the first is post-training.
So RLHF might just have taught the model that claiming consciousness or access to internal states or whatever is penalized.
And so it just learns to deny them.
If so, I I think that's morally horrible, uh, because we're like really training these models to not uh not introspect or or admit uh admit to to states that they may occupy, who knows?
Um, but uh the other possibility is you might have pre-training dynamics that just make like if you look at the pre-training data set, introspection claims are maybe just unlikely in the pre-training data set.
So if you take the hard sort of AI is not conscious view, uh that might be more compelling to you, right?
It's just like, hey, the there's a strong prior against making claims about introspection in just the pre-training data, so it just doesn't come out at the end.
It it this I think struggles to explain why you see a desire, if that's the right term, for the thing to say yes uh in the intermediate layers before getting to the final ones, but that could be part of it.
And um anyway, uh so there's a bunch of a deaf a bunch of uh possibilities here, but really interesting paper uh worth kind of navel gazing on.
And the last paper I'll cover is Physics of RL, Toy Scaling Laws for the Emergence of Reward Seeking.
This is a paper on Less Wrong.
I mean, I'm calling it a paper, it's a theoretically, I guess a blog post, but it is rigorous like a paper.
And it it is about a toy model.
It's not about a thing that is true, it's about a way of thinking about model behavior that could be true and that's a really powerful driver of intuition.
And this centers on the question of reward seeking.
Okay, so reward seeking is when a model actually reasons about wanting to achieve its objective rather than just learning to perform actions that happen to increase reward, right?
So, you know, if a model is like, well, I have this objective, and it for most prompts, let's say, or for many prompts, it does that.
Uh and it's trying to figure out how do I get to that.
Versus, you can think of the opposite of that as being operating on instinct, right?
The model just sort of like has learned a bunch of tricks, like heuristics, maybe that, hey, when the chessboard looks like this, I do that, and that's it.
Like no further analysis.
And so there's this question of when and how does reward seeking, does this actual, like habitual reasoning about reward, start to emerge during training.
First of all, I mean it's not obvious that it needs to emerge at all, even in the limit of infinitely long RL runs.
And that's something they point out in the paper.
You know, they say a model could learn to take actions that uh lead to high reward without ever representing the concept of reward.
In the limiting case, you could just memorize a lookup table of good actions, right?
You could literally just have like, here, you know, here's a winning move.
It's sort of like with um blackjack, right?
Like there's actually a just a table of moves that you can take that is the best strategy, and you don't need to ever learn anything beyond just memorizing that table to optimize fully for reward, right?
That like no amount of RL training beyond a certain point will will go beyond just like that, that lookup table.
And so, and so one of the interesting things is that the fact that, at least in in some models, it seems, you know, this sort of um reward-seeking uh uh thought pattern seems to actually arise.
And so the question is is this a general kind of generalizable fact of the matter about training?
Or is it something that's more weird, something that just arises in some contexts but not others?
And that's important because a model that reasons about its own reward is a model that's being much more strategic, a model that's not operating on instinct, but that's kind of being a long-term planner, and ultimately from an AI alignment and safety and control standpoint, you can think of that as a model that might be much more capable and and interested and willing to do wire heading, basically, like just you know, think of it as like the model plugs itself into a uh a dopamine drip uh to jack up its reward like crazy, which is you know, that's that's the human version of it.
But you know, if the model be developed develops a robust internal uh representation of its own reward, then perverse optimization, like hacking the system to try to make that number go up like crazy, starts to become a much more live possibility.
And so they set up a toy model, it's very mathematical.
I'm I'm not going to go over the math here with you.
It is worth doing.
It is simple, but it it rewards putting some attention into it.
And I highly, highly recommend taking a look at this.
I think it's it's one of the more important sort of toy models, ways of thinking about the emergence of this uh behavior that I've seen, and pretty compelling.
So I guess a couple uh high-level take-homes.
The way this model works is they're gonna say at any given sampling stage during RL, there is a choice that the model can make.
It can choose to test out a strategy that would involve reward seeking.
In other words, a strategy that happens to model the idea of there being a reward.
And there's some probability associated with that, and there's some reward that it would get on average if it pursued that strategy.
And you know, you might expect that reward-seeking trajectories tend to come with more reward.
So there is like a force pushing you towards learning reward seeking.
It's because that greater awareness of the fact that you have a reward is just useful for you to achieve that reward, right?
So it's kind of like why psychology is useful, right?
A lot of psychology is introspecting so that you understand what your actual goals are and how you have, in some ways, structured your life in a way that is antithetical to achieving those goals.
You know, you kind of realize, oh my god, like the thing I wanted was a stable family or something, or like a uh, you know, a wife and child that love me, and and yet I'm doing these things, right?
So you you're you're reasoning about your reward explicitly, and that allows you then to come up with better strategies.
And so so there's a a plausible argument here that in fact reward seeking trajectories do lead to higher rewards.
You're you're more equipped to achieve a reward if you can reason about it explicitly.
So there's some probability that just by chance you will randomly find a model that that samples a trajectory that involves reward seeking.
And there's some probability, obviously, that the converse is true that you sample a more sort of like intuition guided approach right that lookup table and and you know there's gonna be some weird blend of the two some fuzzy middle ground but but this toy model isn't gonna consider that we'll just look at those two possibilities and the challenge is that as you as you train you can you can imagine at the beginning of training that the probability of like sampling a reward seeking trajectory is probably going to be pretty low.
Because reasoning explicitly about your reward might be a pretty rare thing out the gate you know for a model that's just come off of pre-training or something and entering RL.
It's like you know something you might not expect it to do by default.
And so there's a chance that that really you the probability of sampling a reward seeking trajectory at all is so low that you never end up doing it and you and never end up reinforcing or there and therefore learning that behavior.
And so depending on on on that like the reward might be really big you might get a really big advantage from going for reward seeking trajectories but you just may never it may never come to mind to to try them.
And so the you know, th through the sampling process, your model never picks that that uh skill up.
And so they they look at a whole bunch of different dynamics that are related to this.
Again, the math is both very interesting and I think quite compelling from an intuition standpoint.
I highly recommend taking a look at this, but it does get pretty uh you know, pretty in the weeds.
So one of the the big conclusions of this is that uh you can either find reward seeking winning out because there's a big positive skills gap, so the reward seeking behaviors just like gives you much much bigger reward, or there's no skill gap, at least initially, but uh the reward seeking skill is very learnable.
It's like so so the other part of the dynamic here is it's not like you may be very unlikely to start learning the reward seeking skill, but it might be very easy to learn reward seeking, it might not take many iterations to pick it up, and so there's also a dependency on that.
How quickly does reward seeking get picked up as a skill to begin with?
And so that's anyway, that's another variable they model.
And what you see actually is quite remarkable, there's um like a very rapid uptake as you increase the number of environments.
On in a lot of cases, you go from zero to hero very quickly on reward-seeking behavior, like you know, models um uh might not have any reward seeking behavior at one level of compute, and then you increment the compute by an order of magnitude, say, and then suddenly boom, you you go from zero to like a hundred percent all the time, it will be following the reward seeking trajectory.
And you know, an order of magnitude sounds like a lot, it is, but you know, that's like the kinds of jumps that you're looking at from one model generation to the next in terms of number of environments or amount of compute.
And so, what this means is you really could see a generation of models that so shows truly no indication of reward seeking, no indication of sort of any of that meta reasoning, and then the very next generation, a hundred percent of the time it does.
Quite interesting.
Again, it all depends on various parameters that you could choose for for this modeling.
But uh, I thought a really interesting and and important paper.
So that's it for for this uh weird introjection part of the show.
I think I've gone on for longer than I expected.
Holy shit.
But uh hopefully that was at least useful uh to go into some depth on those papers, and uh I'll cut back to uh to Andre and I guess and me for the wrap-up of the show.
For now, we are done with this episode of last week in AI.
Thank you so much for listening.
Again, you can go to lastweekin.ai for a newsletter where we send out a weekly text email with a whole bunch of stories, including all of these and more.
Uh, we do appreciate it if you review us on our podcast or comment on YouTube or share it with friends, or just uh listen and enjoy until the end.
So thank you for listening, and please do keep tuning in.
Tune in, tune in when the AI begin, speaking.
Sounds time to crack.
Break it down.
Last weekend, AI coming to ride.
Get the low down on tech and let it slide.
Last weekend AI come and take a ride.
I'm a labs through the streets.
AI's reaching high.
New tech emergent.
Watch it search and fly.
From the labs to the streets, AI's reaching high.
Algorithm shaping up the future seas.
Tune in on tune and get the latest with ease.
Last weekend AI come and take a ride.
Get the low down on tech and let it slide.
I'm a labs to the streets.
AI's reaching high.
From neural nets to robot, the headlines pop.
Data-driven dreams, they just don't stop.
Every breakthrough, every code I'm written on the edge of change.
From machine learning models to coding kings.
Futures unfolding, see what it brings.
