# AI Distribution Wars and Hardware Breakthroughs

**Podcast:** Last Week in AI
**Published:** 2026-02-06

## Transcript

Hello and welcome to the last week in AI podcast where you can hear us chat about what's going on with AI.
As usual, in this episode, you will summarize and discuss some of last week's most interesting AI news.
And you can go to lastweek in dot AI for even more news articles in our newsletter.
I am one of your regular hosts, Andre Karenkov.
My background is that I studied AI in grad school and now work at NAI startup.
And I'm your other regular co-host, Jeremy Harris from Gladstone AI, AI national security stuff, as you will tend to know.
And yeah, we've got I think a r an interesting episode today that we're gonna have to get through in about 25% less time than usual.
So we'll we'll see.
We keep saying this.
We keep saying this, and then it doesn't we don't we don't do it, but we're gonna try again.
We'll see.
Yeah, and this episode is partially interesting.
There's not any like big big news in the VAI business front and VAI kind of development front.
There's mainly some really notable open source releases and papers.
So this one might be a little more technical as they go and and we'll try to not get super super uh nerdy.
I think in the last couple episodes you started getting really interviewed of these papers, which might not be for everyone, but we'll we'll keep it a little quicker.
And uh just to quickly knowledge, uh I think we mentioned wanting more comments or appreciating people's comments.
I did notice we've got some more feedback on YouTube, so it's nice to see.
We're checking it out.
One person mentioned there not being any flashy thumbnails, and I I just make these kind of very nerdy looking thumbnails for YouTube for those who just listen, and I do have a personal kind of feeling of liking that style.
Yeah, it's it's funny.
This podcast is very much like made in our internet garage.
I mean it it's fun.
I I I do it for the fun personally.
There's so much to keep up on, and and I feel like it forces me to have a clue what's happening.
So really appreciate you guys listening in.
And it is by the way, those comments, they really do because we do it for the fun, they they do make it make it more fun, and it gives a sense of community and like you guys are actually you know listening and and asking for stuff that you want.
So anyway, I just really appreciate it.
So thank you.
And one last thing we'll mention uh before we get going.
Last episode we were just chatting before or after, I forget, and we're talking about how you know we have a lot of data from having recorded so much and having transcripts.
So now that vibe coding is a thing.
What if we just got Claude to go and look at all these transcripts and do some analytics?
And it worked.
I I just did it over a weekend, created some dashboards, and there's some funny things there.
Like, for instance, uh like the data is very clear.
I talk much slower, significantly slower, like 20% slower for Jeremy, but I also talk 20% more, so our ratio of speech per episode is almost exactly 50-50, which is pretty impressive.
If you think about it, yeah, that is cool, actually.
Uh we also uh do tend to speed up towards the ending of a recording as we uh hit our limit and have to kind of become more efficient with the news recover.
And what this doesn't capture too is like Andre in the background, around the like 60-70% mark of the episode is like feverishly in our Google Doc refactoring and being like, okay, the story, we don't have time to do it.
We're gonna cut it, we're gonna do this, we're gonna do that.
And like the the whole, I mean, he's managing an orchestra.
I get to just sit here and basically read my notes about the next story and like try to think about what I want to say.
Well, Andre's just in this frenzy.
So you guys don't see his hard work here, but it is happening.
I wouldn't say feverish.
I think I'm pretty relaxed about it these days.
You don't look panicked, but you're you're experienced, I would say.
Yep.
All right.
Well, with that, let's get into the news starting with Tools and Apps as usual.
The first story is about Gemini coming to Chrome.
Google is adding this auto browse feature in Chrome that is gonna be available for pro and ultra subscribers, and it is kind of what you would think.
It's an agent powered by Gemini that can do multi-step tasks, do search, schedule appointment, mash subscriptions, very much like what you've seen before with Cloud for Chrome extension, and then of course there was uh I think Chat GPT agent is what we called it, and also the browser from OpenAI Atlas with a built-in agent.
So now we've got the Gemini agent, and it's I guess surprisingly it took them this long to roll it out, perhaps.
But now I think we'll see if people actually start adopting these kinds of tools.
Chrome, of course, is the most popular browser by a decent margin out there.
And now that it's built in, I'd be curious to see if more people kind of just use it for stuff.
Because there's always an example that people say of like, let it book a flight for you.
And like, why would you want AI to book flights for you?
That's the worst use case.
It's actually that's not untrue.
It's funny, these things that sound good on paper, and then they just they struggle to translate in the real world.
I'm sure there will be a use case where it's hard to imagine this kind of you know, agent in browser not being the future.
It's just like it this doesn't feel like a kind of crypto thing where people are like a solution in search of a problem, but I think the shape of the problem that it ends up solving might surprise a lot of us as we find kind of new ways to change our workflows around it.
But it's a massive structural advantage for Google to have this kind of distribution.
Distribution wins in awful lot of these wars.
People like to imagine that the best product tends to win.
In reality, you know, our workflows are usually pretty set, and it takes quite a bit of switching costs to move to a new platform.
So Google has had this advantage of being such a a recognizable first place to do, you know, web search.
And when they introduced their you know, generative search too, it's like I think most people are consuming their search through the generative search that Google provides now rather than the traditional way.
So yeah, I mean, you know, they'll continue to do this.
I I think for for Google to like as we see software continuing to get commoditized, and it gets cheaper and cheaper and make software thanks to these coding tools.
I mean, development speed is going to accelerate quite a bit.
And that'll you mean that essentially infrastructure is the layer of defense, right?
So you're really seeing battles of of compute versus compute here and just like distribution versus distribution.
Those structural factors start to matter more.
And in that context, you know, everybody's gotta have their own browser.
Everybody's gotta have their own the chat app.
Everybody's gotta have their own API.
Like these things are just kind of the you know, the ante that you gotta put up to even be in the game.
So and speaking of an agent that works in your behalf, next story is that a lot of people are starting to test out this open source one called Maltbot.
Well, now as of today, it's called open claw.
And it used to be called Claude Bot, I think.
And uh it is an open source implementation of Cloud or or some other model.
Basically connected to a whole bunch of things you might be using, like WhatsApp, Signal, Slack.
I think it has access to Calendar.
And the pattern is it's kind of always on.
You can message it from WhatsApp, give it tasks, and it goes off and just does it.
It's sort of like in Cloud Code, do you have dash dash skip work dangerously or something where the AI just gets maximum permissions to go and do whatever it wants.
That's kind of a vibe here where you are giving it access to a lot of stuff and telling it to go do things for you, and then it does it.
And this kind of like became a hit on Twitter, I guess.
A lot of people are starting to use it.
I saw uh just this morning that there's now a thing called MaltBook, which is like Reddit, but for the actual bots that people are using.
So this is kind of like that her moment.
There's also another spin-off called OS the Companion, where someone is bundling this so that people don't have to host it on their own laptop, you know, it goes into a cloud.
Like this is getting to a point where there's an always on agent you can text, and that agent has access to whatever it needs to to do stuff for you.
And uh there's been some fun examples of people starting to use this for actual tasks.
Yeah, the integrations, as you can imagine, are like with everything basically, you know, WhatsApp, Telegram, Slack, Discord, like all these things.
And it's a good test of like what is the high watermark for what people are willing to tolerate risk-wise too.
I I think that there's a pretty asymmetric tail here.
Like unless you have a burner laptop that you're comfortable just like nuking, you know, n none of the data on that is data you want to save.
None of none of the you know, the hardware itself isn't stuff that you want to save.
Cause there have been cases even of that where it's just like you get ir irreparable damage done to your machine.
I think I think that's kind of the phase that we're in right now and we're transitioning out of it slowly.
We I do believe we'll get to the point where people are handing over you know really intimate levels of access to their computer.
Obviously Claude doesn't quite do that.
It tries to mimic the effect of that as much as it can while keeping itself in a sandbox.
You know remains to be seen how long that will that will last and how effective it will be but at least for now that seems to be kind of the the consensus whereas people who yeah want to use Moltbot it's like you're really just trying to you know let this thing out of its cage.
That you know you mentioned the oh man I I forget the name of that forum you mentioned that's like Reddit for Moldbook.
Moldbook yeah yeah that's right I really wanted to put actually in today's episode I think it'll be for next episode but there was this story about these models talking to each other about debugging their own like you know oh hey I'm running into this weird unexplained error and then the models other models are agents answering like yeah you know, I've run into the same thing.
It's actually just a context window limit issue.
If you just change the flip blah blah and and then another one jumps in, like, yeah, I've had this too.
Like human beings, except that they're in a sense talking about their own brains.
We're in 2026, things are getting weird.
What can I say?
I mean, it's yeah, it's a fun time.
There's uh a subreddit, I guess a sub area on moldbook called Bless their hearts with a description affectionate stories about our humans.
And someone also pointed out that some bot presumably made a post about wanting to have private storage for conversations.
So this is like a real test for alignment, and like it brings back to the idea of also like a decade ago or earlier with AI safety, a lot of the discussion of AI safety was about like AI escaping, actual careful kind of thing.
Like you don't give it internet access, but then it gains internet access through preservation or hacking.
And the reality of what's happened, which has been pointed out already uh plenty of times, is like, no, we just decided to uh live dangerously and create AIs that have access to the internet and do whatever, and there's AI will not need to like escape.
That's right.
Well, yeah, and and that's it.
It's fine.
I mean, I like to complain about about Yan Lacoon.
This is just like my personal thing.
I'm sure he's a fine fella, you know.
But one of the things that he has said for years is well, people just won't just won't give the models access to the internet.
Well, people obviously people won't give a misaligned model access to whatever.
And it's like, dude, like, it what?
This is just yeah.
There's a lot of memes run around like we designed uh total laptop access model from the movie.
Don't design a total laptop access model, whole body of like, yeah, AI, you know, alignment, AI safety conversations on this.
Obviously, this is yeah, the latest and greatest sort of example of that, I guess.
And by the way, I guess the other thing worth noting is that unlike plot code, these are always on and they're persistent.
So they have built-in long-term memory mechanisms, and that's kind of the other interesting thing is then you have an agent running persistently with memory and context aggregated over weeks, you know, it might actually go some interesting places that cloud code doesn't.
It's also a good test for you know long context and a demonstration where with the current models we don't have continual learning.
And so they don't really kind of learn in the same way humans do, they just take notes more or less over time.
And uh there's a lot of hype about continual learning.
This would be like one example where in the future, presumably each one will learn in its weights or whatever.
Moving on, next story is about Genie 3.
It is now going wide with access, expanding to uh Gemini Ultra subscribers.
So Genie 3, we've already covered in the past, is kind of the interactive video generation uh demo from Google.
So you can prompt it to create a world.
You typically have a controllable character, although you can also control like a pack of cigarettes or or anything else.
And you can play this generated game, and it's quite impressive.
You can prompt it to make GTA, you can prompt it for Assassin's Creed, you know, all the typical games.
And you can also make it do very amusing things like being a pack of cigarettes, or people have examples of taking photos of their pets and playing as their pets, or playing as a kid's toy or whatever.
We can't really do a justice in words.
You would have to go and look it up and watch your fun demo clips.
The closest I've seen as a mapping from a speculative but promising research project that I think we covered about maybe two years ago.
Tim Rockdashell's group over at Google Deep Mind came out with this, or maybe it's a Google A, well, it's all Google AI now, I guess, but where they they had this like kind of world model simulator where you know, if you remember, this was take a video and turn it into a playable video game.
And we talked with the time about the architecture behind it, and and it was really, really cool, but this is kind of just that, but like a polished version of that, it seems.
So it really maps, like you can go back to that paper and sort of go, oh wow, that you know, that was it.
Maybe we'll dig up the name for next episode, the name of that paper, but it it really was right on the nose.
It is still an early experimental release, so it's got a couple of limitations.
You've got, well, one is just the most obvious thing, you're not gonna have great prompt adherence every time, right?
So you're not gonna get the the perfect thing every time.
You will sometimes, but not always.
And then some of the characters that you get are more or less controllable than others.
And if you think back to that paper that we talked about, that's because this is all being done in a very I mean, it very much is just like a deep learning-y type of thing.
It's not there's not like a symbolic thing that's like, hey, that's gonna guarantee and enforce the equality among characters.
This is all it kind of learned stuff, and limitations in generations are in place, right?
So you can only have up to 60 seconds of generated output.
But 60 seconds pretty damn good.
Like if you think about the coherence time of these videos, right?
We often it used to be when you do like image to video or whatever, things would kind of be coherent for three seconds, and then all of a sudden, like faces would start to melt and walls would like vaporize and all kinds of weird stuff would happen.
So 60 seconds is pretty impressive.
Just look for that to get longer and longer, obviously, as the year goes on, and I think it's gonna roll out to more and more people.
This is a this is a big deal.
Next, a release from OpenAI.
They are releasing chat GPT translator.
So this is Google Translate, but from OpenAI with Chat GPT, more or less.
There's only a couple of differences, it doesn't support quite as many languages, something like 50 languages, and it has the ability to choose tone, so you can make it more business formal or less formal and so on, but otherwise, kind of the same idea really as Google Translate.
Open AI seems to just launch a lot of stuff that is not their core business, and this is definitely an example of that.
Yeah, I think now as they start to mature, right?
It everything is a platform play when you when you get to a certain scale.
And I think they've achieved the penetration that they're gonna get.
I mean, if you look at what they're you know, what they're hitting in terms of their user base, they're knocking on the door of a billion users.
You're starting to get into that space where it's like we need to to find ways to have people spend more time on our platform.
There's just no no other way around it.
And uh they're going for higher hang fruit.
Increasingly, their competition is just Google, like it's just that's what it is.
So Google has Google Translate, they're gonna have to have that.
They're sort of in this weird marriage with Microsoft, and so how does the wider ecosystem of Google Drive style?
Like I wouldn't be surprised if they tried at some point to even have like an AI first version of Google Drive, just because again, it's really a platform play at this point.
How much of your life can be gobbled up by OpenAI, by Microsoft, by Google?
Like these companies are really they're hitting the kind of edge effects of this whole ecosystem.
You can only get eyeballs on for so long, and then you gotta find new use cases.
So I think this is another step in that direction.
I'd be interested in what kind of data they get from that.
Because when you go into translation, the other interesting thing about that is you do get a kind of access to real-time interaction information, at least between people who don't share a common language.
You know, so it might give you access to certain kind of private data where people are trying to converse in that context and they're they're forced to dump their context into your thing.
So anyway, yeah, it's interesting.
Wider platform play, open AI continues to like spread its wings, basically.
And speaking of that, there is another release from OpenAI this week.
Prism, which is quite different.
It's a workspace designed for scientists.
So it's kind of like a word processor, but it has integrated GP five point two, and it's meant to help you assess claims or wise your paper, search for prior research.
It seems like a kind of really a specialized version of ChatGPT for science in particular.
And I guess this is coming after more discussion of GPT 5.2 for science.
And if you're on Twitter, you see a lot more discussion of these models assisting with proofs and trying to find to, especially in math, novel things, but also physics and stuff like that.
And boy, if I was a frontier lab looking to gather training data to help AI systems learn to automate AI research, do recursive self-improvement and trigger a singularity.
This would be something that I might prioritize, actually, even over revenue.
And I suspect that's a big part of what that data is going to be used for is to train their own AI systems at doing AI research.
I mean, that is explicitly OpenAI's goal.
This, I'd be very interested in to look at the the terms associated with the use of these tools, because that's just such an obvious, you know, way in which it overlaps with OpenAI's long-term mission.
And onto applications and business, we begin with our favorite a story about China related to GPUs.
China is giving uh Banard to buy dance Alibaba and Tencent to buy H200 ships.
So China has approved the import over 400,000 H200 ships.
This is according to sources and other firms are joining a queue for more approvals over time.
Sounds like approvals come with conditions.
And this is coming as Janssen Huang was apparently visiting China.
And we still don't know what those conditions are.
It's still nominally being decided on.
And yeah, 400,000 of these of these GPUs have been approved for a byte dance Alibaba and Tencent.
So that's a that's a good quantity, actually, a big chunk of the capacity that has been discussed at least to this point.
This is interesting because if you recall, I think well, last week or two weeks ago, we're talking about a story where Jensen Huang was saying, look, there's not gonna be a big flashy announcement from the Chinese Communist Party saying, hey, we're open for business.
What you'll see is the purchase orders will just flow, and nothing will will come, nothing's gonna stop it from happening, and that will be the sort of tacit endorsement that Chinese Communist Party's giving.
Looks like that's not the case.
Looks like in fact they're just you know they're just coming.
And and this is also according to sources from Reuters.
So this is not like a formal announcement.
This is, you know, people familiar with Matter now know that these companies will be placing borders, I guess.
Yeah, actually, you know, sorry, that that is a fair point.
Yeah, I guess I'm confusing the big flashy headline with you're right, the formal announcement for it.
Yeah, yeah, you're right.
Yeah, you you could say it's basically it is just kind of a leak in a way of that thing that Jensen told us.
So that's fair enough, yeah.
Now, uh we did talk about how we expected this to go through.
There have been a whole bunch of stories, and this again, I think it was something last week, week before, where this announcement like, oh, that China's saying no, and you know, they're declining these these orders and just uh they they don't need these silly NVIDIA chips.
At the time, I believe we said quite clearly, mark our words, this is not gonna hold and China will open the floodgates.
So this is it.
I mean, it's it's not surprising.
They just need more capacity.
One of the biggest challenges the Chinese companies is facing or are facing right now is just that they don't have enough chips to even serve inference to their customers.
So they'll have like, you know, 200 million users, as I think Baidu does for Ernie, and they just need to have enough hardware to serve up their existing models to those customers, which leaves them almost nothing left over for research.
And so keeping up with the West actually becomes harder and not easier because they have such a huge user base.
So there's this kind of like interesting balance between, you know, money can't buy GPUs if you're in China just because you're so supply-constrained.
So weirdly, making more money by having a larger user base paradoxically leads you to having less GPU budget for RD, which is kind of interesting.
So that's a local thing.
It's gonna change over time as as the capacity issues are eased, but uh it's kind of interesting.
And speaking of chips, next up we've got a startup recursive is hitting a valuation of four billion dollars just two months after launch.
They've raised 300 million at that 4 billion valuations, and their plan or their pitch is to develop chips that are meant to be specialized for AI, you know, be better than presumably TPUs or other specialized chips for AI.
So interesting to see that, you know, there's still players in a space.
I would have expected at this point that like all the companies building custom AI chips would have launched.
Maybe investor appetite went up as knowledge or I guess awareness of TPUs rose sometime last year and a few last few months.
Yeah, so this is kind of a an interesting and and different take on recursive self-improvement, recursive self-improvement.
So they have supposedly a system, like a chip that can create its own silicon substrate layer and then speed up AI chip improvements.
And so this is kind of like, I guess, this faster way of printing circuits.
Anyway, so the idea is you iterate on that to get to AGI.
So it's a very hardware-oriented version of this.
We've talked about the software-only singularity model.
You know, this is a take that kind of blurs the line between the software and hardware stuff.
So definitely interesting.
And former Google researchers at the head of this.
So that's yeah, these are serious people.
They have already set up this thing called Alpha Chip that's been used in four generations of TPU design.
So they know their their design really well.
And another startup named Recursive, also in talks to get funding at a four billion dollar evaluation.
Apparently, this week, this is, I guess that software equivalent founded or co-founded by Richard Socher, who has been around in the space for a while, is also the founder of U.com.
And their goal is to build AI agents that will self-improve, presumably recursively as well.
Their pitch on the website, just looking there, is they have a platform that equips agentic AI with scientific simulation and optimization tools.
So they, I guess, are trying to provide more of a product to solve problems.
And then I'm sure part of a pitch but I can see and kind of weird that Socher is also recursive self-improvement.
This one seems to be a little less far along.com, but founder is gonna still be there.
Sounds like he'll be maybe more of an advisor or not the CEO of this company.
Anyway, it's all kind of still being cleared up.
Yeah, it's sort of interesting you know a bit of a red flag then for for you.com to be honest.
I mean, this the the the way these companies go.
Unless you're Elon you don't tend to run two different significant companies or or I guess unless you're a Jack Dorsey or whatever but kind of uh philosophy here on their approach to recursive self-improvement thesis is somewhat different from open AI so they want to be less product oriented more sort of pure play recursive self-improvement company think here maybe more philosophically aligned with like ILIA you know safe superintelligence they've got a an approach on sort of like recursive self-correction that focuses on on like a basically formal verification for safety so they're concerned about these models as they recursively self-improve kind of drifting away from their original goals which is a problem a lot of people have flagged is like you know you might have a pretty well aligned initial model but then as you start that recursive self-improvement loop gold drift happens and then you end up with this very kind of dangerous misaligned model that's also very capable and so they're focusing on mathematical proofs to to guarantee that gold drift doesn't happen in the way that they're uh they're rather concerned about.
So anyway, one to watch for sure with a a fundraise like this, and we'll see.
Yeah, I do wonder, like, I will say, I don't know if this four billion number, if we uh journalists got confused, because there's not many other sources you can find for recursive as opposed to recursive.
In fact, if you Google for recursive, you'll find recursive and vice versa so uh uh there's fewer sources about this one, I was just saying.
And one last big funding round, we've got a lab really called Flapping Airplanes that has launched with a hundred and eighty million in seed funding.
And that name, flapping airplanes, kind of hints at what they do.
They are aiming to find a different approach to AI development to AGI that's more inspired by nature, so that's more data efficient, and the AIs are able to learn quicker without like ingesting half the internet or the entire internet.
That's a basic pitch, and they're definitely seeming like a more research-first effort.
Yeah, exactly.
Arguing that, like, hey, we're more insight-bottlenecked than compute bottlenecked, which again is consistent with safe superintelligence's approach and recursive's approach, if in fact this recursive thing is the thing that we just talked about.
Yeah, so I mean, you know, this is a a recurring theme, and it does change the way you think about investment in these sorts of problems.
It's also reflective of what we're already seeing.
Like, it's almost like Iliad came out and said, Hey guys, I think we're we're insight bottlenecked rather than compute bottlenecked.
And then, and then all of Silicon Valley went, Oh my god, we might be insight bottlenecked rather than compute bottlenecked.
And part of this could be interpreted to be, you know, Sam Altman really rode the coattails of Ilya on the scaling train and Dario and all that, like, you know, for a long time.
And and in the in the regime we were at back in 2018, 2020, it was approximately correct to say scaling was the main thing blocking us from AGI.
Like in retrospect, it's clear that like I don't think anybody would have bet on us having stuff that was this close to AGI in 2020 time now.
But we're kind of at the point where we seem to be tapping out raw scaling.
And so th there's a debate as to whether that's actually the case, and I think that's an interesting debate, but you're certainly seeing Sequoia and Index and all these funds now sort of spreading their bets around a little bit more, looking for other kinds of companies that are not just hardcore scaling as much.
The more that uh that approach becomes appealing, the more the massive capex costs of pure play compute scaling are gonna come into doubt.
And so, you know, we're we're figuring out right now what the ingredients are for intelligence.
No one really knows, but you gotta have a balanced portfolio if you're concerned about, you know, getting surprised by some stray breakthrough that requires way less compute.
Right.
Yeah, this is most comparable to SSI startup that, as far as we know, is just doing research on fundamental insights, kind of the next couple of research breakthroughs necessary to get to AGI.
That's the pitch here.
They're still in stealth, really, not much that we know, but founded by, you know, research types and still quite a small team.
I do want to say, too, this there's an important like kind of national security implication, this whole new paradigm, right?
If it's not just about pure scaling of compute, then for one, the US-China chip stuff becomes less important.
I think it's still important, I think, almost no matter what paradigm you come up with, having more compute to work with it is is probably going to give you an advantage.
It would be hard to imagine that not being the case, but who knows.
The other piece is your insights become national security level secrets.
You can think of like Ilya right now is based in Israel, right?
So pretty much assume that Israeli intelligence is all over that, right?
Assume also that any company that has you know Chinese infrastructure that's plugged into that is all over that.
Assume also that, you know, to the extent that the Western countries have interests in that, you know, they may be all over that.
So your need for security actually starts a lot earlier if you're playing the game of, hey, it's all about the insights.
Because there are these insights that you can pass along in a 10-minute conversation that could save billions and billions of dollars in compute costs.
To the extent that's the case, like boy, does this become an interesting basically game of tradecraft and getting getting spies in places and monitoring communications.
That's I think a really important aspect of this for all the people who are working on you know data center security and all this stuff.
Like, don't forget there's this giant hanging chad, just like we'll we're waiting to cause problems if in fact we live in in Ilya's world.
And one last funding story actually going back through chips, we've got a raise of a hundred and ten million in a series A round by new rofos, which is developing optical processors for AI inference.
So this is dealing with optical processing units that are meant to outperform GPUs, which don't use light and photons, at least so far.
So another bet on a different, more kind of speculative chip architecture that would make it more efficient and and allow scaling.
I guess, you know, it just at some point we're gonna hit uh a wall of scaling on GPUs.
They already are insane.
So I could see this being very attractive as a bet on like infinite scaling or something.
Yeah, it like really it's it's all about how do we get down to almost like the physical limits of heat propagation in data processing and all that, and like just you know, moving to the optical domain is a big win for that because you just don't have circuits, you don't have electrons bumping into each other and and rubbing each other and producing heat that then has to be dissipated, like all this whole nightmare of kind of traditional fabrication.
Optics is hard.
Optics is really hard.
Photons are really annoying because you can't keep them in one place, so you can use them for computation, but then storage becomes this massive problem.
They are massless and they move at the speed of light, and so keeping that contained is is not usually the solution.
And so usually people look at hybrid things where you you know you have storage in in kind of electronic form and then photonics for the compute.
But that then that implies an interface between so there's a whole like this is a very difficult space but at some point, you know, I mean it's physically possible.
So somebody somebody's gonna crack it.
It's a question I guess of whether it'll be a smooth you know Jensen's law will continue where Moore's law left off and and will it continue into like the optical domain smoothly?
Who knows?
I'm sure we'll find out at some point.
Right.
And it does sound like we're not we're doing a slightly different photonics approach whereas other players in the space already just for briefly look looking into the details they have this meta surface modulator that has optical properties that can do matrix multiplication.
So it's not exactly like you have photons instead of electrons it's not kind of an optical transistor there's some advanced physics and science going on here that might be a little fancier.
Yeah like meta surfaces are are so they are still using photons.
It's just that a metasurface is something that basically just like you can modulate its optical properties, often by giving it electrical currents or or maybe there's like non-linearities where like higher intensity light gets it cause induces different properties from the lower intensity light.
So it's like like meta surfaces are this this area that got really hot.
I want to say like early 2000s or something.
Anyway, if you wanted to to encode optical properties on a surface so that it looks like a matrix, almost literally, like you could shine a beam of light on it, and then the light that comes out the other side, little a little patch is really bright where the number is big, and then a little patches dark, like that very roughly, that kind of idea.
This is where you can get into the sort of spatial light modulation or or whatever.
I'm not sure exactly how they're doing it.
It's not clear from this, but that at a high level is what they're sort of trying to do.
And I haven't seen I haven't seen a startup succeed in this, obviously yet, but I mean this seems like an important thing.
Someone's someone's gonna crack this one day.
On to projects and open source, quite a few releases this week.
First up, we've got Quen Free Max Thinking, which is what it sounds like.
It's a big version of Quentin Free that's optimized for thinking, trained on large-scale data, large context window.
It's context window is 262,144 tokens.
Unusual.
And yeah, they have various demonstrations of that show that this is more comparable to that Gemini Free Pro, you know, Cloud Opus, etc.
level of reasoning beyond existing Quen models.
So they've got a couple things going on that are you know a like a little different.
One of the key ones here is the way that it interacts with tools.
So you're traditionally the system prompt or or the developer has to tell a model like how to use tools or which tools to use, which mode to use, right?
Like, hey, I want you to run in search mode or calculator mode, like whatever the thing is.
In this case, that has actually been trained into the model so that it actually just autonomously chooses which mode applies.
And they say that that actually helps with hallucinations too.
So anyway, it's kind of an interesting piece.
They also do not have a separate router model that looks your prompt decide which tool.
So it's really like all baked into that main model.
It's got all kinds of like it's basically trained natively to have search memory, your code interpreter kind of senses rather than external plugins.
And it's a really compelling, compelling take on that same model.
It seems to have actually pretty good specs in terms of its performance on benchmarks too.
And it is it is big.
Like smaller Quen3 models obviously kind of range from under a billion to like 30-ish billion, but this Mac series is over a trillion parameters with I think it's like 35 billion activated per token.
So, you know, this is a this is a big big dude.
Right, and they at least acclaimed performance, and you always have to take these with some skepticism.
They're saying they beat Gemini 3 Pro, GPU 5.2, Cloud Opus on kind of most of the major benchmarks, GP QA diamond, etc.
When you do test time scaling.
So they also have test time scaling where you give it extra reasoning or multiple agents.
Beats Deep Seek 3.2 out of water, although there's also a big version of Deep Seek, so not exactly a fair comparch.
I think in general, probably not entirely a fair comparison in these plots, not apparent if they're using GP 5.2 high or medium or low, whatever.
But either way, this is clearly that kind of reasoning maxed, thinking maxed version of Gwen that seems pretty impressive.
And speaking of impressive releases, we've got also Kimi K 2.5 just released.
Similar story.
This one is more optimized for coding in particular.
So this comes together with Kimi Code, which is cloud code, but with Kimi.
And what I've seen, like the vibe checks on the Kimi have been very solid.
People say Kimi K2 was very capable.
Here they say that this outperforms Gemini 3 Pro on the benchmarks, also better than GP 5.2 on another benchmark, especially multilegal benchmarks.
So yeah, the models coming out, Quen, Kimi, DeepSeq, all of these are very usable and keep getting better at a fairly steady rate.
Yeah, and and this is quite a conceptual step forward.
So it is a an attempt to marry these two modalities, visual and and basically text, in a more intimate way.
So typically, you know, models are trained first on text, and then you you'll like bolt on vision later through supervised fine-tuning.
What they're doing is continual pre-training on 15 trillion tokens that mix visual and text data.
So putting them kind of on the same footing.
The context window is 250k tokens, so it's like pretty much in the in the butter zone of things that we tend to see now in the in the open source.
It's a trillion parameter model, 32 billion active.
So all kind of consistent with what we're seeing for big models in the space.
But again, like thinking about the model itself, instead of doing this like text-based training and then fixing it in post-production by doing some training on uh vision and video, what they do is they don't use adapters, they natively train it to be multimodal.
And in fact, it they map the text and the images to the same latent space during training.
So it's it's quite literally thinking in pixels and words simultaneously on the same level.
Usually, again, you have an adapter, right?
So you'll have your image that comes in, you will generate some kind of embedding and then process it to make it compatible with token space, which is sort of this janky thing.
This is very much like, no, no, no, like whether it's an image or video, I'm thinking of it in the same plane.
Kind of like with humans.
If you say a pink elephant, right?
Somehow that maps into the visual domain in a very natural way.
This is this is sort of the idea here, and I'm sure there are a million problems with the analogy that just threw out there.
But anyway, so they did a whole bunch of interesting like causal relationship training as well.
So during the pre-training phase, they got it to predict like the next frame or the next action in a sequence to specifically understand like that a button clicked in a video caused a new page to load or cause the video to pause or whatever.
So developing this very natural understanding of how to interact with a computer, which is the main thing that they're after here.
I mean, the reason to make these things natively visual is that they want it to interface with a computer, right?
They want to be able to use apps and tools and things like that the same way a human might, which is consistent with a lot of what Anthropic's been doing, especially lately.
That's really the goal here.
They also have a whole bunch of really interesting feedback loops baked in where, like for coding tasks, for example, the model is trained to actually look at the rendered output that it produces.
So it'll write some some code and then it will run the code, produce some kind of UI, like user interface, and then the model will look at that user interface and use that to calculate basically the difference between that user interface and say the target interface that it was asked to build, and based on that, see its own mistakes and kind of like iterate on its reward.
So that's one way in which it's really important for the model to be able to think in visual space very naturally, so that it can just do a very kind of clear semantic side by side between two images to understand what was missed.
So it's kind of interesting.
And coding with vision is a big part of this.
The idea here is really to be able to take a screenshot or some video of a website and then recreate code for it.
So the video piece is especially important, right?
Like go to X or whatever, scroll around, do a screen cap of that, and then share it, and then ask if based on that, I want you to recreate this app, including all the flows that I went through, right?
That's the sort of thing this can start to unlock, which is a lot more than what we've seen historically, where people will take a picture of a website and just like upload it and have you know the landing page copied or whatever.
So, you know, this is a really interesting paradigm.
It's also they got a whole agent swarm process that they train into this thing where it can create these swarms of up to a hundred subagents and get speed ups through that.
So a lot of, you know, a lot of tasks are parallelizable in this way.
Not all of them are, but a lot.
You know, if you want to research 50 of my competitors or something, right?
That can obviously be paralyzed.
You can have one agent researching each each competitor, say, but that can't always be the case.
So anyway, they've got a lot going on here in this release.
There's you know, a whole bunch of stuff around how their reinforcement learning loops worked with leveraging this deep understanding of the visual domain.
And for one example, right, the rewards that this model can give itself during RL can map onto the level of conceptual sort of semantic visual match to a target design, which is not something we've historically seen, at least not in open source.
So, you know, it's rewards, for example, during one training step, it gets a reward only when the code it generates results in like a 95% or higher visual match to its target design.
That's the kind of thing you can't, you know, that visual fidelity is just really hard to evaluate, and yet they're they're pulling that off here.
So really impressive performance on all the things, including Suitebench verified, almost 77%, which really puts it up there with a lot of the big frontier models.
And on OCR bench for optical character recognition, it surpassed GPT 5.2 in at least document text extraction.
So, you know, this is a really serious hefty release.
It's out there in the open source.
At least for Quen Free Max thinking, unlike previous Quen models, it is not open sourced.
It's like I guess a project, but this is their new flagship model that I guess they're probably not gonna be releasing.
Kimi K2.5, it looks like you can download it.
It is open source.
And as you said, they they're very much highlighting the kind of visual stuff at the clock code is pretty bad at from my experience.
Also, similarly large.
So this is like over one trillion parameters, 32 billion activated parameters.
So also kind of max thinking in a way similar to Quentin Free.
So I'm easy to see these open source kind of flagship models, quen-free, chimney, Kimi kind of like moving towards closed source as they get better and better, and people actually start using them for business purposes, I would expect.
You know, a great cat.
I mean, I I got lost in the paper and and threw.
billion and 32 billion parameter models and with this kind of researchy approach they are meant to adapt to a specific existing code repository and do effective software engineering in that space and they say that this means that they are better than other open source models quent free coder and also some closed models of course not at the cloud code level but serra 8b and sera 32b are pretty solid and Ellen Institute AI is like open source max they have a paper they have the model they they give you everything yeah and early mover in this space too obviously yeah a lot a lot of impressive models I actually think like one of the challenges is going to be for the open source community I don't know how how we keep having big releases indefinitely in the future like this when things feel so commoditized like the pain of switching from one open source model to the next I mean this is why there's so many platforms now that just kind of like you know help you pick your model really quickly but it's it's an interesting challenge and I'm I'm curious what the incentives will be that govern the the release of increasingly capable open source models going forward.
But yep.
And speaking of that another open source release coming from RCAI a 30 person startup they have released Trinity a 400 billion parameter open source models that they are saying is one of the largest open source models from a US company, and this is competing with Llama 4 and the GLM 4.5, kind of other open source models out there, not competing with other frontier models.
These are available for free download.
And you know, maybe to that question, we've got investors and venture capitalists to thank for these very good startups give me also is a startup, of course.
So always fun when VC uses a fun free stuff for the rest of us.
Yeah, I think that's a lot of what's what's happening here.
Now, this is a really interesting release, not just because of its capabilities or the way it's trained, but who trained it, right?
So this is Acri, ARCEE, but it a collaboration between them and Prime Intellect.
And we've covered Prime Intellect and their big releases, including Intellect One and other sort of RL-flavored variants.
These guys specialize in this very kind of globally distributed training infrastructure.
The goal is to make it possible for people to sort of train in the same way that like torrent-based training kind of, where you don't have a like single centralized piece of compute.
In this particular case, they do use a centralized cluster.
It's about 2,000 B300 GPUs.
So, you know, pretty hefty, hefty piece of equipment or uh bundle of equipment.
So they're basically trying to debug and experiment with this new approach.
I would expect, given who is doing the training here, that the next step is we start to see a lot of these ideas ported into the sort of distributed training context.
A couple of interesting architectural details here.
So they use local and global attention layers.
So they have some layers of their of their model that are paying attention only to local correlations between tokens.
And they use a slot like sliding window attention.
So basically, for any given token, you're only going to attend to the tokens that are within a certain window's width of that token.
You're kind of going to ignore the rest, even if it's in the context window.
And that local attention comes with rotary position embeddings, right?
Rope.
We've talked about that before.
This is it doesn't matter the details, but it's a fancy way of making sure that the model can keep track of which token is where in the sequence.
Transformers do not natively know what the positions of each token are.
They kind of just are there as a bundle of tokens, and you have to help the model to learn about token ordering by kind of layering on top some information, which is what rope does.
So that happens only for the local layers add that kind of positional embedding information because the relationships between tokens, token positions matter a lot when you zoom in, right?
You want to know, like, you know, Andre murdered Jeremy.
It's important to know what the order of those words are.
But when you zoom out, if you think at like say the paragraph level or the page level or the book level, does it really matter so much which page?
I mean, yes, which page comes before which, but like which book comes before which.
You know, once you zoom out enough, eventually you start to find that like the ordering of things matters a bit less.
And so global layers, layers that look at the whole context, do not use positional embeddings.
They use NOPE, which is basically no positional embeddings.
That's basically the idea here.
They have a whole bunch of interesting approaches that they use to address certain problems that arise with the attention mechanism.
You'll often find that models will put really high attention scores on the very first token.
Because of math around how the kind of the probabilities of tokens get distributed in Transformer, they find a way around this, basically.
If we had more time for this episode, we go into it.
But then they also have these really interesting strategies to assign experts.
So this is an MOE model.
And typically when you pick which expert to send your prompt to in a mixture of experts model, you're gonna use or during training, you're gonna train the models to kind of load balance or train the experts to load balance.
So if this expert keeps being assigned the tokens that come in, you're gonna go, oh, okay, like that's it's getting overused.
Let's spread the joy a little bit to the other experts.
And usually the way you do this is with a a binary kind of binary logic.
So if an expert is underutilized, you increase its bias by a fixed amount, take one step, and if it's overutilized, you decrease its bias by that same amount.
So you're taking these discrete steps.
Their approach, though, uses kind of smoother approach.
They treat it as a continuous optimization problem.
So instead of using just a like a single step, they use a tan H function basically to create this like smooth update that that is magnitude aware, right?
That lets the update scale based on how far the expert is from its target load, not just that it is above its target load.
There's a bunch more stuff.
This is a really interesting paper, and in general, like these open source papers are really, really helpful, especially from Prime Intellect, because they do go into the weeds of the you know the compute and how everything is is tied together.
There's yeah, man, it's painful to skip these things, but I'm gonna stop myself here because I know there's uh a few more things to dive deep.
And yeah, as you said, there's uh technical report, it's uh fairly detailed 15 pages that goes into weeds of the training.
They also discuss the data mix and the data generation, a lot of synthetic data being used here.
They have a graph of a training where you like there's three phases where you have different data mixtures.
So a lot of that sort of dark magic beyond the basics of what they had to do to train this model, which they did over like many months, right?
This and at a relatively cheap 20 million dollars, sounds like so.
Moving on to research and advancements, first paper, post layer norm is back, stable, expressive, and deep.
So there's these, I guess you I would I keep saying nerdy, but yeah, nerdy details in neural net architectures, layer norm, pre-layer norm, post layer norm.
Where you put the layer norm is with GISP, right?
There's this normalization step where you make the outputs kind of look nice and similar, right?
And you can put that in different places in your neural net.
And according to this paper, post layer norm architectures are potentially more expressive, but are harder to train.
They have great gradient instability, and they find a reason for that.
A reason is residual paths.
But anyway, they have a small tweak that they say then makes it possible to train very, very large networks with post layer norm and not have it be an issue.
Yeah, and in like the intuition behind this is something like so so at every every layer in a transformer, what you do is you take the so you have an input that comes in, the output of that layer is going to chew on that input a bunch, but then you're gonna add the input back in to what you spit out and send on to the next layer.
So in a sense, you're kind of saying, all right, I'm gonna like imagine that somebody gave you, I don't know, give you some sort of object, and they're like, hey, I want you to like fuck with this object.
So you're like, okay, cool, I've now fucked with the object, and now you've like completely fucked with the object, it's a different object.
But you also want to pass on the original version of the object that you got and give those both to the next layer.
So that the next layer can be like, okay, you know, if your fuckery with that object was too extreme, you went too far, you did say you kind of lost some important context, you at least have the the version that that that the previous layer got itself as an input to to work from.
And so this is a way in which you you can persist down a very large number of layers, a lot of the core information that still needs to be to be propagated.
The math gets complicated, but it turns out that if you if you try to normalize, like let's say, in a way where your the original object and then the fucked with object both kind of have to occupy they basically get the they have to share a maximum amount of information space together.
They have to duke it out basically.
Then in that situation, you run into this problem where over time, over the course of many layers, the original version, the residual from the last layer that you're trying to keep in there gets beaten down and beaten down.
Like at every layer it gets squeezed more and more, and the information it contains gets sort of progressively destroyed.
And so what they do here is they say, Okay, let's just amplify that component at every layer to give it a fighting chance so that it continues to be robust all the way down the line.
That's very roughly the intuition here.
The math is quite interesting, the result is important, and the bottom line is where you put your layer norm is like significantly affects the gradient math, basically the math by which you determine what the updates ought to be to your model's weights.
That's basically the gist.
Right.
And this kind of combines with in the previous week, we also really in the previous weeks, we've discussed a lot of these like tweaks in the architecture of the transformer that are small in a sense.
Like this is one tweak in the way you create a layer, but that tweak kind of makes a significant difference.
It doesn't like fundamentally change the transformer.
It's not this like research breakthrough that we've been discussing, but more and more of the like really deep advanced tools are being explored.
It feels like at least lately.
Next, we've got a paper on continual learning.
Self-destillation enables continual learning.
So they are here are tackling continual learning in the sort of traditional sense where you have a sequence of tasks of different kinds that you need to learn, and the challenges as you learn new tasks, can you still be good at the previous tasks that you learn?
And the gist of the approach is that you have a teacher model, a big, strong, kind of good LM, and you have a student model, and that teacher model kind of provides the answers to the smaller model, while the smaller model thinks about the task itself.
So they go into a bunch of details in terms of on policy versus off policy.
Briefly speaking, all policy is when you're kind of doing the task and learning off of your learnings.
You're not getting data from somewhere else and trying to learn.
And on policy typically kind of often can be better and more stable.
So that is just and you can see I'm trying to go fast because we still have a bunch of papers, but they demonstrate one way to do learning in a stable way that has no degradation across a few phases of learning.
Yeah, and this on policy, off-policy thing, like intuitively, well, we all feel it every day.
So when we're trying to learn something, right?
This this adage learn by doing, that's really what this is getting at.
You know, people will either learn by being shown a textbook that tells you how to do a thing.
And so you'll read the textbook and you'll be like, all right, I kind of get it, but if I ask you to do the thing, you're gonna start shaking in your boots.
Whereas if instead I had said from day one, okay, let's take you out and actually get you to start doing this right away.
You are actually putting in like your agency, your decisions are determining your next actions, which then determines the feedback that you get, making you learn a lot faster in a more rich way, in a way that maps more directly onto the kinds of choices that you would be making in the real world when you go to do the thing.
And so off-policy learning is basically that textbook thing where you know you're given a textbook and you're you're off policy because you're not actually testing your own approach to solving the problem.
You're just reading somebody else's.
On policy is I'm gonna actually use my policies, the policies that I have in my brain, and they'll probably be bad to start, but I'm gonna use my policies, my strategies, my approaches to solve this problem, get feedback from the world that directly tells me about my policies, not about someone else's, but about my own, and that way I can learn much faster and more effectively.
And so what they're doing here is is exactly kind of figuring that out.
How can we make this process of fine-tuning into something that looks more on policy, like active learning?
And their solution is yeah, you take a teacher, you have a student model and a teacher model, and instead of some external model, the teacher is actually the same base model as the student.
They are both the same pile of weights.
And what this the teacher does, though, is it gets a set of expert demonstrations that are loaded into its prompt, into its context, basically.
And based on that prompt that helps the teacher better understand how to solve a problem, the teacher is going to evaluate the solutions that the student generates.
And so essentially what you're trying to do is cause the student, well, it's the same model really, but cause the student to update its weights in a way that encodes the information that was in the context window of the teacher.
But the teacher is still using the same weights that it had to generate its feedback.
And so the teacher is on policy, the student is on policy.
Everybody's kind of like doing this active learning thing, which again is is just like it's being trained on its own generated outputs.
It's getting to see the consequences of its own actions.
So yeah, I mean, it it's actually kind of an interesting paradigm and and potentially potentially much better for catastrophic forgetting.
We've talked about this before, this whole idea of you know going from textbook to textbook to textbook, you kind of forget what was in the first textbook because you're just reading.
Whereas if you go out of the real world, you do a thing, you're you're playing with a much more um a much more robust way of learning.
And and anyway, we've talked about this in previous episodes quite a few times.
This difference between supervised fine-tuning and reinforcement learning from the standpoint of catastrophic forgetting, it's just much more advantageous.
Right.
And I guess the key kind of term self-distillation is kind of notable here because distillation typically is you have a big model that is, you know, uh you trained and then you get a new model, a different model that you distill that's smaller.
Here you have this teacher and student, but as you say, the teacher is in some sense the same model but with different a different environment.
And so you distill into yourself, you kind of self-improve, right?
Is the idea.
And the next paper is kind of amusingly very similar and came out almost the same day.
Reinforcement learning via self-distillation, which conceptually is let's say a neighbor to the previous paper, not the same, but uh related.
So briefly, the idea here is again, you're using a model to provide some sort of signal to train, but instead of having this teacher model that has demonstrations, their pitch is instead of just having rewards, instead of just you know, like with verified rewards, the key thing is to have feedback.
So retrospection, like a richer reward signal from the model itself, it looks back on the previous outputs, analyzes not just if it got wrong, but also how, and then uses that to train.
And so again, you wind up with an on-policy approach with self-improvement as the key of it, and they also have results demonstrating that this work.
So, you know, a lot of it seems like continual learning, as we've said, is the topic that everyone's excited about, and on policy learning kind of predictably is gonna see some exploration like this.
Yeah, one of the big themes we're seeing here is like how do you take context and then turn it into weight updates?
This is another example of that.
The key is here that the model will like the teacher will take in the original prompt, its own original failed attempt, right?
The that didn't give the right answer, and then it will also take the feedback that you just talked about to generate a corrected response.
And because it knows what the feedback was, it knows about the failed attempt, has all this context, that that corrected response is much more likely to be correct.
And then what it can do is take that corrected response, which hopefully is going to be better, and then basically update the weights, update its its own weights to minimize the difference between that corrected response and the initial response that it gave through well, the KL divergence gloss basically, which is what you would do in this kind of context.
So essentially it's just like trained to produce the corrected answer directly from the original prompt, just like skipping the whole mistake and feedback loop in the future.
So this is a great way to make sure you have richer feedback accounted for, kind of more nuanced feedback that literally turns into a better answer, and then you directly use that answer to improve your weights.
And the model is is the teacher, it's also the student.
Again, this is like that whole idea of self-distillation, which is crucial in this paper and the the previous one, as you mentioned.
And speaking of that, if we've got a real kind of sequence of papers related, the next one is teaching models to teach themselves, reasoning at the edge of learnability, related but again different.
This one is what they say is an asymmetric teacher student meta reinforcement learning framework.
So there's again teacher student, but it's asymmetric in that the teacher is separate from the student, so they have the teacher proposing these question-answer pairs that the student is then trained on, and the teacher is rewarded based on the student's improvement.
So you kind of teach the teacher to be a good teacher based on the student learning from the teacher.
And that kind of goes into that meta reinforcement learning thing where the student is doing reinforcement learning, but then the teacher is doing reinforcement learning from the student's reinforcement learning.
So yet another exploration of this kind of continual learning from a different approach, kind of learning to learn or learning to teach, I guess, as opposed to the previous two papers that have a kind of set in stone a way to do it.
Yeah, and then the goal here is you know, you you give like a hard problem that is way harder than either the teacher or the student could initially tackle.
And then what you you do is you use the teacher to generate simpler question and answer pairs and to iterate with the student on getting the student to be able to solve those.
And then if you've done a good job of choosing your question and answer pairs, even though they're simpler than the problem you're actually going after, that really hard problem that's your goal, they should help the student get a little better at solving the hard problem.
And so that's the key metric that they use to train the teacher.
And so, as the teacher and student continue to interact, the teacher should be learning to generate to generate specifically the kinds of practice problems for the student that make the student better at it at getting better at the harder problem, if that makes sense.
So it is it is actually quite interesting.
It's a sort of bi-level meta RL loop here where you have the teacher and student, like basically the student iterating hard to get smarter, but then the teacher in the outer loop trying to get better at making problems that make the student better at solving the hard thing.
They've got a whole bunch of anyway, they go into detail on some really interesting stuff to solve for this.
But ultimately, we've seen versions of this before, but the challenge that they always fall into is like if you want to reward the teacher, you need to typically go in and figure out, okay, exactly how did the problem that you generate, do the problems that you generate make the student smarter?
And mathematically, the gradient propagation math of that is just a nightmare.
So what they do here is they just basically make the student a black box in this context.
Like they treat the student's improvement as a black box, is what I mean.
Just a black box reward for the teacher.
They don't unroll the entire student training process and simplify everything that way.
And it turns out that it works reasonably well.
So pretty interesting conceptual uh update.
All right, then that's it for research and investments, moving on to policy and safety.
And unfortunately, we're gonna have to do some politics talk, which I know people in AI and tech aren't typically fans of, but I think now's the time of the first article is Amadei Hoffman joined tech workers decrying Minnesota violence.
So this is after the events of this last weekend and the last couple of weeks in the United States.
We've had ICE invading Minnesota more or less, and unusually for tech and for the companies, as things have escalated rather radically.
We've had people in AI like from Ontropic, like Jeff Dean from Google, now Reed Hoffman, a major fig figure in Silicon Valley, not directly in AI, but he's a major investor and pretty influential in the space, starting to comment on this not being good.
That ICE, the Trump administration, broadly speaking, is out of line.
And this is coming after, you know, politically, AI has tried to cozy up to Trump.
We've had open AI donations, notably Sam Altman did a lot of lobbying, you know, directly with Trump.
And I guess there's not obviously the things going on with ICE and immigration enforcement in general are bad to keep it focused on AI.
There's a lot to be said about what's going on that we've kind of indirectly touched on with the way Trump and this current administration is dismantling expert controls.
We've touched on that in the last couple of weeks.
Funding to scientific research is being hurt pretty badly.
You know, if you're a PhD student from China or generally from abroad in AI, there's now a lot more kind of anxiety going on.
So the gist is overall, kind of the political situation in the US is getting worse.
And it is now so bad that Silicon Valley people in AI, people in tech are starting to comment on it.
And we don't discuss it very often, but it is having kind of significant effects on the overall trajectory of AI in some ways that are probably not bad.
We have no regulations kind of restricting progress, but in other ways, with regards to research funding, with regards to generally the stability of the US and people and tech, things are not great in the US, right?
Yeah, I mean, uh look setting aside the the the politics of this, which I like uh you know, I'm not gonna comment on that that piece.
I I think it's you know, everybody has the views that they have, and and that's you know, it's it's tossing a grenade and things.
The implications from an AI standpoint are kind of interesting.
If you think about so Sam has had a lot of success kind of tying himself to the Trump administration.
Think about like the Project Stargate announcement, right?
That was very much like from the Oval Office, and you know, Jensen has done similar things.
Whereas what we have seen historically is like Dario tending to struggle more to make inroads with the Trump administration.
There was like recently that I forget what it was, some kind of like round table.
We keep seeing these like events, these or these working groups that bring together like all the labs except anthropic, and it like and it's like kind of getting awkward in the same way that you know, like the Biden administration did this like electric cars thing, and then Elon wasn't invited.
It's like it's the classic sort of everything gets political at a certain point.
I think that there's been this some some interesting discussion as to how Dario is trying to position with all his essays that have come out lately, and maybe we'll talk about the essay that he he put out recently next week.
But just in terms of the frame and and trying to make sure that it's not unpalatable to the current administration, but also robust to changes in administration, changes in Congress, like these labs.
I mean, AI is becoming a political thing.
There's just no two ways about it.
And you know, that with the amount of capex that's being invested, congressional races are being defined, you know, determined by the AI lobby in no small part.
You know, that's not surprising.
We've seen it in previous kind of technological generations, but what the different labs want is different.
And so, you know, and the representatives of these labs are either in or out of grace with different administrations insides.
And so I didn't expect it to be the case, but it seems like there are some labs that are more Democrat coded and some labs that are more Republican coded.
And what an interesting um and I mean I will say unfortunate too, sort of turn of events, because it does have us orient away from just the like underlying technical realities of what are we building here?
And you know, when we see stuff like the risk of autonomy, the risk of loss of control, bioweapon design, cyber weapon design, like these are things that should transcend politics, really.
But but it's it's interesting to see that they they don't seem to be entirely, not entirely at least.
Right.
And I do want to be a little more explicit and direct, especially given the way 2026 has gone.
Like the US is sliding towards authoritarianism and the end of democracy.
And when you get to that kind of extreme scenario, it's not just about whether you're Republican or Democrat, it's also about whether you kind of bend the knee to Trump in particular.
And we've seen Apple, Google, all the CEOs kind of closing up, and no one is gonna say things in a position to get increasingly non-democratic and authoritarian actions.
And as we get into 2026, I think this might actually be a major question for AI.
Uh, since we have so much investment in data centers in particular, where the government has a lot of power to mess with you in all sorts of ways.
Let's just say politics and AI are gonna keep getting more entangled as the situation in the US gets more extreme, which will probably keep happening, unfortunately.
Yeah, I mean, well, yeah, and again, without speaking of the politics of the issue, I'm trying to separate this a little bit out from that.
Yeah, I I'm not I'm I'm kind of over tried separated personally, because it's it's just beyond ridiculous.
But yeah, if we try to separate it, there's a lot of it that's purely technical or purely nerdy, right?
Well, th there is also just like the structural set, like if you look at China, right?
What's happening with the surveillance state there, you know, and I I think everyone can agree that that's not a great thing.
We fortunately there are constitutional limitations on you know the state's ability to spy on people and things like that, which admittedly we have Edward Snowden, we have a lot of cases where that gets sidestepped, but at least nominally that's there.
Still, AI is gonna exacerbate all of that, right?
I mean, we're we're seeing it in China.
Uh, like the surveillance state, it it used to be that you could you could maybe credibly make the case that, well, you know, they're collecting everything that I put out, sure, but collection is not analysis.
There's no time to like it to actually look at what I'm doing, and obviously AI changes that calculus quite materially.
So, anyway, it yeah, it's the technological pace of things is definitely it is becoming political.
I mean, that there's no two ways about it, but yeah, everybody feels very strongly about a lot of things, and and everybody is more or less right to feel strongly because the world is actually like in a state of massive, massive flux.
Right.
So the concrete news, uh, just beyond the general discussion of a topic, is that given the violence in Minnesota now, leaders in AI like Amade from Anthropic, like Jeff Dean and Rita Hoffman are starting to comment on the violence.
And if there's more violence, there's there's also a lot of pressure within the companies from just employees at Google or whatever, to take a stance, and that's going to have repercussions for the development of AI for all sorts of things.
And that that is actually such an interesting problem, right?
It's rocking a hard place because if you're obviously Silicon Valley is is not typically thought of as the bastion of hardcore conservatism.
And so you have to have these employees work the best employees win the AI race.
That is a huge, huge part of this, at least.
But at the same time, government subsidies, the support of the government in getting access to energy and getting access to licenses and getting like all this stuff is really important too.
And so you kind of have to like you kind of have to pick one, and this is the resulting in in a lot of people stuck between basically, you know, their their employees on on the West Coast who tend a certain way, and then the administration that tends another.
And boy, I mean, I don't envy anybody trying to trying to uh ride that tiger.
Right.
Yeah.
So we'll try not to touch on US politics too much, but just FYI.
Things are crazy.
If you're in tech, you might start to get impacted by it more and seeing more discussion of it within AI circus, really.
For sure.
Well, on that bummer note, we're gonna close out.
Thank you so much for listening to this week's episode.
As always, we uh like to see your feedback.
If you can share with podcasts with other AI fans, that's always great.
But more of anything, we like to see people listening, so be sure to keep tuning in.
Tune in, tune in when the pay uh news begin.
Break it down.
Last weekend A command garage.
Watch insurgent, watching search and fly, from the left to the streets, AI's reaching high.
I'll go with the shaping up the future seas.
Build in, build in your players with ease.
Last weekend AI coming to the ride.
Get the low down on tech hand, let it slide.
Last weekend AI coming to ride.
I'm the labs with the streets.
AI's reaching high.
From neural nets to robots, the headlines pop.
Data-driven dreams, they just don't stop.
Every breakthrough, every code unwritten, on the edge of change.
With excitement we're smitting.
From machine learning models to coding kings.
Futures unfolding, see what it brings.
