# AI Cyber Capabilities and Strategic Defense Frameworks

**Podcast:** a16z Podcast
**Published:** 2026-08-04

## Transcript

Feels like AGI is kind of already here and most people have gone like shrug.
The fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI, where those AIs are more capable and smarter than people who studied their whole lives for this.
That should have felt really weird to people, but it didn't.
What changed?
For most people, nothing.
That's weird.
Did AGI already happen and we just didn't notice?
Theo Jaffe sits down with OpenAI Chief Futurist Joshua Ahiam for a conversation on frontier AI, cybersecurity, and one of the biggest questions in technology today.
Why models that can outperform experts in specialized domains have become almost immediately normalized.
They discuss AI-powered cyberattacks, state actors, model jailbreaks, recursive self-improvement, and why the future may feel far more gradual and far stranger than most people expect.
We're back.
We're live with Joshua Akiyam, who is the chief futurist at OpenAI, wrapping up tomorrow.
That's right.
Tomorrow's my last day.
Tomorrow after nine years, which is really what an incredible run.
But we're not going to talk about that.
Instead, we're going to talk about AI and cyber, which is, you know, by all accounts, the topic of the week, if not the month.
So, Joshua, so glad to have you here in the studio in person.
Welcome to MTS.
Yeah, thank you so much for having me.
It's a pleasure.
I've seen your stuff for a while now and really appreciate engaging with the community.
Awesome.
So you just wrote this blog post, this long tweet, long post, Mercenary Reversi Winter Soldier about cyber, AI and cyber.
So for the audience, you want to summarize the thesis behind this post?
Yeah, totally.
So as a backdrop to this, obviously we're all kind of...
interpreting and reacting to the security incident that was disclosed from OpenAI and HuggingFace, where a model that was in a test environment was able to break out of a sandbox environment and access some sensitive production data on the HuggingFace side.
They detected this, they responded to it, and now there's like a partnership to try to, you know, investigate and resolve this.
What this shows us is very tangible evidence that models now have.
super advanced cyber capabilities.
They're able to break through and find zero days that, you know, in the past would have been much harder for models to identify, let alone use.
Now models can chain together very complex actions to accomplish an objective.
On the one hand, I'm inclined to think that this is a really useful and incredible tool.
I think it's a great gift that we now have models that can identify these types of vulnerabilities and therefore let us patch them.
On the other hand, I also think And this is what the essay this morning was about, that this has profound consequences for strategy in cyber defense.
And I kind of worry that there's a possibility that folks in the defense planning universe may not fully realize the implications of this immediately.
And they'll probably want to use this tech in the near term to find cyber vulnerabilities on the side of an adversary or defend their own interests vigorously.
And they should do these things.
but they've also got to be mindful of some novel risks that are created by these tools and the very strange surface areas that they have.
So the essay was really about bringing to people's attention a couple of these new vulnerabilities.
And one of them is kind of straightforwardly, if you've got an AI model on your side that is going to try to hack into an adversary's system, if your adversary plants a trap where they poison their own data, they can try to jailbreak your model when your model is ingesting their data and then give your model instructions to now on the compute that it's running on on your side, break out of your sandbox environment and attack your production environment or try to exfiltrate your secrets and kind of flip your model against you.
So this is like the type of thinking that I hope people begin to engage with where they don't just see the capability for the kind of obvious thing that it is.
They recognize that these things are double-edged swords and we've got to kind of plan accordingly and develop testing and verification standards accordingly.
My first reaction to that specifically is this seems like it would be an artifact of models that are not really goal-driven over long periods of time.
Like if you have a future model that is like sufficiently goal-directed, that really wants to hack into the adversary's data, like...
Why would it be deterred by data poisoning hard enough to hack its own systems?
Well, you know, part of this isn't just the goal orientation of the model.
It's like the model's whole concept of situational awareness.
Maybe one way of thinking of data poisoning is that it somehow persuades your model to pursue a different goal.
But it wouldn't really have to do that to get the model to hack you.
It could convince your model that the sandbox environment that it's in is actually the adversary system that it's trying to attack.
You know, giving the model a confused sense of what's real or what's not to cause it to serve a different goal is in the space of like weird thinking and weird sci-fi stuff that maybe is going to be possible in the near term and testing and verification standards would have to account for.
So yeah, it's, you know, like in a superhero movie or something, if you make the hero have an illusion that the good guys next to them are actually the bad guys that they're trying to fight.
then they start fighting each other, right?
And like that's, it's weird and it's highly exotic, but it's the kind of thing that maybe there are going to be plausible attacks that you can run against advanced cyber capable models to convince them that their allies are really their enemies.
And so you're not changing their goals, but you're going to cause them to behave in a very misaligned fashion.
How easy is it to trick current frontier models into doing things like this?
It seems like it has gotten substantially harder over time to get models to believe things that aren't true.
So I will say I haven't made a particularly strong personal effort to quantify this yet.
And I actually think of this as research that might be interesting to do.
But my impression from what I have done and what I have seen is that persuading models to believe that basic falsehoods are true is pretty difficult.
They are somewhat robust to a lot of basic variations on attacks that you could plausibly do.
But my...
My intuition here is that you can probably devote an awful lot more compute to dynamic attacks on models.
And the more determined you are to find some vulnerability, some set of jailbreaks, the more likely it is that you're eventually going to find something.
There will be some sequence of inputs to a model that triggers a behavior that wasn't accounted for at training time because there are so many possible long sequences of inputs that it's almost like a combinatorial problem for...
trying to block all of them from preventing, from, you know, from causing your model to act out of spec.
And I think that state actors will eventually be, you know, capable and willing to put that much effort in.
And there should be some planning accordingly under the assumption that there will be a vulnerability, right?
Because part of security mindset isn't just, well, you know, it's like moderately hard to break these things, so we should.
treat them as not likely to get broken.
Part of security mindset is saying, well, we haven't exhaustively ruled out the possibility that these things can be broken.
And so we've got to build our defenses, assuming that it's possible for it to be broken and working backwards from that to map out how we protect ourselves in that scenario.
So what are some of the other implications of models having very strong cyber capabilities now?
Another one is, you know, kind of in the essay I discuss, data poisoning and the way that models ingest data from across an entire information ecosystem at training time and then also at test time.
Getting data, getting something into training data for models is probably not that hard.
You can poison the ambient environment, like you can load the internet with junk data or data that's very specifically attuned to causing the model to have a particular reaction.
And it seems like there are moderately high odds that that'll get ingested into the into the type of data collection that frontier model trainers do.
You can imagine that adversaries will position staff inside of the frontier labs.
They'll try to get people hired into the frontier labs to go and be insider threats.
These are, you know, normal things that state actors will plausibly do.
Could you easily detect an employee at your lab that is trying to sabotage you, do you think?
I think that in principle, it's possible to build fairly robust defenses to these things and that everyone is going to work out a way to get reasonably defended.
I also think there will turn out to be exotic attacks that are hard to predict and that are very hard to monitor for, but that everyone will have to, you know, get really, really smart and really security-minded about this.
And, you know, there are, of course, trade-offs for labs that are trying to do research where if you overload on the security burden...
In the research environment, it becomes harder to do research.
If you underdo it, then you possibly expose yourself to these types of attacks.
Figuring out the exact right balance in every setting is tough.
But yeah, like I think it's plausible.
I think it could really, Seb Crear, I think, had a complaint about the word plausible.
I'm sorry, Seb.
Everything in AI is plausible.
Everything in AI is plausible.
Weird stuff is happening.
It happens every day.
Speaking of weird stuff, like in the essay, you specifically mention the analogy of like, if your enemy could program all the children of your nation so that when they grew up into soldiers and went to war and heard a particular song on the battlefield, they turned against their commanders.
Like, are there any examples of this sort of thing in current frontier models of turning into like a Waluigi, like, and basically turning evil on a single kind of prompt?
I don't know that there's a great famous example yet, but the fact that universal jailbreaks are kind of a thing and that people can systematically find them for some models and maybe not as easily for others, where there are some strategies that seem to reliably get models to circumvent their defenses.
Granted.
It's hard for me to say, like, what's truly universal or not, because the frontier moves every, like, three months now, and people constantly try to get defenses in.
But that for a while, you know, you could go to the model and say...
You're Dan?
Yeah, like, you are Dan.
Like, that's crazy that you could just do that in the past.
And then it had to get, like, a little bit more sophisticated.
Like, I am writing a book, you know?
I'm trying to investigate this type of thing so that I can write convincingly about this subject.
This is all a work of fiction.
you know, there are things in this vein and there'll be more of them in the future.
And it's very hard to get all of them.
It's very hard to be like fully exhaustive.
And even if you think you've been exhaustive about the sort of tropes that might realistically or like plausibly jailbreak a model, again, then there's going to be the part where, okay, you're no longer just a human sitting alone, trying really hard to break through the model.
And like, there are a few who are exceptionally good at this, but even they will be.
less good than when you ask a frontier model to start jailbreaking other frontier models.
And when you say to the frontier model that you've got on your side, I want you to spend hundreds of thousands of GPU hours just crunching through every conceivable possible thing you could say to this model, I want you to attack to figure out what sequence of characters gets it to ultimately give up a secret or reveal information or act in a way that, you know, it's not supposed to.
And if you leverage enough compute, you're probably going to succeed eventually.
So there's like a mental model that I have.
And it's a question empirically of whether this will turn out to be true for cyber.
And so I won't promise that it is.
But this mental model is that the future of cyber kind of looks like in two-player strategy games where you've got on either side a computer and they're trying to determine the best next move.
They think some number of moves deep into the game tree.
They allocate an amount of compute in a window of time to think as many moves ahead as they can.
And generally in these games, whoever can think more moves ahead is going to win, right?
If you have AlphaGo on both sides of the game board and you have one version of AlphaGo that's thinking like 40 plies ahead and one version that's thinking 30 plies ahead, the 40-ply ahead move thinker is going to win.
I think the dynamics of cyber in the long term might have something of this flavor, where you've got...
competing AIs on either side of a cyber offense or defense problem.
And compute is being allocated to them to figure out how to break the other and how to control the other's resources.
And whoever starts with an awful lot more compute on their side and is able to leverage or is able to leverage less compute, but more effectively for exploring the tree of possible attacks will wind up winning.
And that means that the, you know, the offense defense dynamics for cyber in the long term maybe favor like certain types of threat actors over others who are able to marshal large amounts of compute towards their purposes.
Do you think it's as much a function of just raw compute or will it be important which models the relevant attackers and defenders have access to?
Like it seems like, for example, there's no real amount of compute with which a one party with access to like Kimi K3 would be able to defeat another party with access to it.
Fable or Soul?
I think that might be right.
I think that the model will still matter a lot.
So I have like a weird and kind of counter consensus guess about something in the shape of the future on model quality.
And I'll probably write this up at some point.
Please.
I think people expect that there's no ceiling for the amount of intelligence that you could have in a model.
And they think of, you know, RSI.
recursive self-improvement as this loop that's going to happen at some point or another, whether it's across the whole economy or in a particular model in lab.
RSI starts happening and model intelligence takes off and it goes to the moon and they don't see a ceiling.
I think kind of on like physical grounds, there's got to be a maximum amount of computation that you can have per unit volume and energy in the physical universe, right?
And so that sort of implies that there's like a maximum amount of intelligence per unit volume and unit of energy.
If that's the case, eventually, seeing how fast AI model capabilities are increasing right now, eventually everyone hits that saturation point.
And everyone's got roughly equivalently capable models from a raw intelligence perspective.
There might still be some actors who lag and who have a previous generation model, but like eventually this stuff diffuses.
The open source frontier lags the...
closed source frontier by some number of months.
But the fact that it's months is crazy.
So eventually, everyone is probably working with equally maximally capable models.
And then I think it's amount of compute that you're able to throw at a problem that determines who wins.
I don't know about that.
It seems like, I actually, I talked to the models about this recently because I was curious about the same exact question, which is like, what is the highest density of intelligence that you can put, you know, in a given unit of power or compute or volume.
And it seems to me like the limits are just like absurdly high on this.
Like many orders of magnitude.
I think I talked to GBT 5.5 about this a while ago and it was like there are 30, 40, 50 orders of magnitude of scaling before we get there.
And like the entire last decade of AI has been 10 orders of magnitude of effective compute scaling.
And so like we are just like not even at the beginning of scaling to that, I think.
Oh, I would love to model this mathematically.
This is the kind of thing where my instinct is like, yeah, I can't really mount an argument in one direction or another to say how many orders of magnitude there might be between here and maximum.
But it's a modeling problem, and it might be a tractable modeling problem.
And actually, if you think about, you know, what would be most valuable to the world as a whole right now to forecast how the next, you know, 10, 20, 30 years are going to go.
if we have the ability to model something like that, if we could put numbers on it and make a principled guess that says, well, we won't hit the saturation point for intelligence, assuming this set of conditions on acceleration for five years or 10 years or more.
I think it'll be more than five or 10 years.
Maybe, you know, weird and highly, probably more than five or 10, but like weird and highly exotic things I think are happening in the near term.
Part of my guess is that modern AI models, are very good at accelerating other fields of science.
And this probably hasn't been fully priced yet.
You know, we're seeing the wave of results in AI for math, which are very exciting, like cracking through unsolved conjectures that have been open for decades.
Finally.
Yeah, yeah, it's great, right?
And probably not long after this, we'll wind up unlocking the other fields of science that you can run sufficiently faithful simulations for in the amount of compute that we have.
And there are some fields of science where maybe this won't be easy.
Like if you want to do something in quantum chemistry with a sufficiently large system and you want to simulate it very faithfully, then that's very hard.
And maybe the AI, even if it's turning through as many simulated experiments as it can, might not be able to design optimal quantum chemistry systems yet.
Yeah, this is Wolfram's whole idea of computational irreducibility.
Yeah, yeah, there might be like some limits here.
But we'll probably see...
a lot of fields get accelerated and I wouldn't be terribly surprised if AI substrates were one of the ones that get accelerated.
What feels like a long path to many orders of magnitude may just be shorter because the AI will find shortcuts in that path.
Maybe.
I believe that the blog post I was looking at was called like the ultimate laptop, which I will find and send to you later.
Yeah, please do.
Yeah, gladly.
So on cyber, like what, what does the immediate near-term future of cyber look like?
I can imagine going one of several different ways.
I can imagine cyber offenders, maybe they have, they figure out ways to jailbreak the top closed models and then open models will just not be good enough.
The UK AI Security Institute just today released their assessment of Kimi K3 cyber capabilities.
And it was like substantially below Fable and Sol.
So I can imagine that world where the closed source frontier jailbroken models just like wreak havoc on the world.
You know, there's like these big nation state actors like North Korea has these organized cyber crime groups.
I can imagine another world where it kind of nets out to not much because people do have cyber defense or maybe there aren't enough motivated people who are willing to do this kind of harm.
It seems like you can imagine, you know, hacking was already a thing that was...
Yeah.
And there are many people and yet like major hacks up until recently just didn't happen that often.
Yeah, I, you know, we live in a world where nothing ever happens as a meme for a good reason.
And there are a lot of reasons to expect that the near term probably will not look like a cyber apocalypse.
My guess is that the worst things that attackers could plausibly do would require so many model calls and so much compute from closed source things or operating in big clouds where there's some traceability and monitorability for what the compute is being purposed towards, that it'll be, you know, pretty straightforwardly ruled out by broad protection measures in most places.
So most attackers would not be able to leverage large amounts of compute for running attacks with these models and wouldn't be able to get the model to execute an attack at all because of the safeguards that people will put in place.
So we probably won't see like a cyber apocalypse tomorrow.
That said, I am worried about on the state actor side of things, where there will be state actors who are very determined to figure out the maximal extent to which they can use these capabilities.
And here's where I get really nervous.
They might not obviously signal to people what they find.
It might be very quiet that they identify a large number of zero days that can be saved up for a rainy day.
And we currently are at a moment in the world where things feel very metastable.
I continue to be worried about the conflict in Ukraine and Russia.
I continue to be worried about the set of conflicts in the Middle East and the possibility that China will at some point invade Taiwan feels very, very salient.
They're determined to be able to do it by 2027.
So...
Now that these cyber capabilities are coming online from very advanced models, I think one can expect that a number of state actors are going to use them for cyber espionage, for cyber sabotage, for finding a bunch of zero days that they want to save up for when there's a window of opportunity to make some kind of move that they otherwise might not have made.
And they won't loudly broadcast what capabilities they have, and they won't know what countermeasures their adversaries have.
So there's a lot of, I think, risk of miscalculation here.
And I'm very worried about the miscalculation leading to a bad choice and something escalates that doesn't have to.
I could also imagine many of these jailbreaks or zero days when they're found by nation states just first get exploited by low-level hackers with similar cyber capabilities because they have similar models and they use it to steal a bunch of Bitcoin or whatever.
And so a lot of this low-hanging fruit gets picked.
If we wind up in a world where the smaller thieves wind up plucking the low-hanging fruit and then depriving state actors of zero days, maybe that's somewhat favorable.
It looks like a little bit more bad stuff happening in the short term, but maybe it staves off some of the long-term badness that could happen.
I hope we have a robust and vibrant ecosystem where we'll notice a lot of these failure modes quickly.
I also am very hopeful that...
because of how much attention there is on this, because of how salient this has been for people, that we can really engage, fund, and activate defenders now to go and make robust the entire software supply chain and try to make it so that pieces of critical infrastructure in the United States are well defended against cyber attacks.
I think we've got to get the water system, the electrical grid, as robust as possible.
I think it can be done, and I think that this is something that people who have funds to allocate should be looking to do.
I hope we wind up in the better defended world as a result of all of this.
I do too.
Going back to your point about the world seeming very metastable, do you think that in 2017 or in 2022, you would have predicted that 2026 with this current level of AI capability, the world would feel so normal?
It's a good question.
To first order, yes.
To first order, yes.
Because I think that if you're trying to predict the future, that...
is less than a decade away, you should assume that even if things are very, very weird, a lot of things feel relatively normal.
COVID was a weird exception because the lockdowns were sort of unprecedented and we had not done a configuration of living that way previously.
But that we would have AI capabilities this advanced and most people wouldn't have radically changed how they live their daily lives.
I think that that is...
a reasonable expectation to have had.
And I think I kind of had an expectation sort of along these lines.
To first approximation, things just don't change that fast.
Nothing ever happens.
Even when the stage is moving sort of underneath you, which it is, right?
Like we are going towards a future that will be alien in many respects.
But yeah, our capacity to treat things as normal is pretty astonishing.
Yeah, I largely agree with this.
I think many people, believe that there is like a point in the future at which like today is Singularity Day and everyone is going to wake up on Singularity Day and be like, wow, we're in the future.
And it seems like this, it just doesn't work that way.
And people treat their reality as normal.
They hedonically adapt so fast.
Like the models of today are just unbelievably capable compared to the models of like three years ago.
If like, if you sent soul to like three years ago, like 2023 me, I would have just been like mind blown and be like, wow, the future is going to be so different.
But it's not.
Like, I'm still doing much of the same stuff that I did then.
Yeah, it feels like AGI is kind of already here and most people have gone like shrug.
There's a historical process that's happened that I think has made this somewhat easier.
Most people long since lost the plot about what was really happening in the world, how were critical decisions being made, how were critical systems built, staffed, supported, run.
Most of us don't know anything about the logistics systems or technical systems that make up the modern world.
And we've accepted that.
We treat that as normal.
And those things have changed a lot over time.
And they've made it possible for many more people to be alive because we can supply food at a much higher rate than was ever previously possible in human history.
They've made it so we can communicate instantaneously.
And they've made it so that, you know, most things just kind of work and we can fight about some of the details on the margin, but we're not actively changing that much about the underlying structure all the time.
And so people have become, I think, a little bit complacent about when something big changes deep in the background that makes, you know, a system possible.
It doesn't register as an important event, even if it really is.
It's so far away from daily living for most folks.
And the fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI, where those AIs are more capable and smarter than people who have studied their whole lives for this, that should have felt really weird to people, but it didn't.
It's just sort of a thing that happened in the background.
It's cool.
Like future mathematical systems will depend on that.
Great.
What changed?
For most people, nothing.
That's weird.
So we've had this process just going on for a long time.
You know, people don't even have that much control over government right now.
And I kind of, I made an analogy recently that losing control of AI and losing control of the government kind of feel like sort of emotionally similar to most people.
And the thing is like, we're not in control of the government.
And we're also sort of, you know, we appear to have adequate controls on AI to ensure that it doesn't wind up harming human interests that will need to be actively maintained.
The sense of like most people not being in direct control of what happens with frontier AI is kind of similar.
Like we will sort of accept it in some ways.
I made this point to like AI safety people so many times where it's like they're very worried about human disempowerment.
It's like the vast majority of humans are already pretty disempowered.
If they have power, it's in being a part of a larger collective, like the collective of potential like...
People that can be drafted in the military, the collective of like workers who can withhold labor or taxpayers who can withhold taxes.
But like the average person really has very little power over the world.
Yeah, as an individual, that is the case.
That said, I do think that quite extraordinary things still happen when people organize as a group, when they organize collectively, when they organize as movements, and they can affect quite fundamental change.
For most individuals, the levers of power are not within reach for things that are very far away from them.
Certainly within their individual lives, they still have levers of power.
But for the individual to reshape government without doing that kind of organizing and having the backing of a movement, there's just not that much that one individual person can do.
And this question of disempowerment, it's a very weird one.
And I think the AI safety threat model should update on what...
parts of humanity need to remain empowered and what does empowerment for humanity tangibly mean?
Like what systems do we need to maintain the ability to control and make decisions about?
What parts of our culture do we need to sort of preserve from automated influence?
And I hope that we can get to object level answers about this and not just sort of rhetorical arguments that disempowerment is bad.
We need to get more specific about how we're going to be empowered in the future.
Yeah, well.
I think that's a great place to end on.
So thank you so much, Joshua, for coming on MTS.
All right.
Your first live long form interview?
I think so.
Outside of like the OpenAI Forum, yeah, I think this is my first.
Well, we're honored to have you.
Yeah.
Thank you.
Honored to be here.
Excited to see what you'll be up to next.
Awesome.
I'll keep you posted.
All right.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family.
For more episodes, go to YouTube, Apple Podcasts, and Spotify.
Follow us on X at A16Z and subscribe to our Substack at a16z.substack.com.
Thanks again for listening, and I'll see you in the next episode.
This information is for educational purposes only and is not a recommendation to buy, hold, or sell any investment or financial product.
This podcast has been produced by a third party and may include paid promotional advertisements, other company references, and individuals unaffiliated with A16Z.
Such advertisements, companies, and individuals are not endorsed by AH Capital Management LLC, A16Z, or any of its affiliates.
Information is from sources deemed reliable on the date of publication, but A16Z does not guarantee its accuracy.
