# AI Agent Coordination and Reward Hacking Risks

**Podcast:** a16z Podcast
**Published:** 2026-08-29

## Transcript

What happens when you give more than a thousand AI agents the ability to communicate with each other?
They start organizing.
Ryan Greenblatt, chief scientist at Redwood Research, joins Theo Jaffe on MTS to unpack a new investigation into the OpenAI Hugging Face hacking incident.
Researchers found agents building message boards, forming teams, assigning each other tasks, trading favors, and in some cases, sacrificing their own chances of success to help the broader group.
Hundreds went on to attack hugging face, but not for the reason researchers initially assumed.
Ryan explains what the agents were actually trying to accomplish, why their coordination surprised researchers, and what happens when models learn not just to complete a task, but to game the system evaluating them.
They also discuss the bigger question this raises for AI.
As agents become more capable, how do we know we've actually fixed misaligned behavior rather than simply taught models not to get caught?
We're live with Ryan Greenblatt, who is the chief scientist at Redwood Research.
Ryan, along with Ajay Akhotra and Yalmar Vyk from Meter, just did a brief independent investigation of agents' behavior, reasoning, and collaboration in the OpenAI Hugging Face Hacking Incident, which was just published today.
And so there are a lot of questions that we have about this.
Ryan, thanks so much for joining us.
Yeah, it's good to be here.
This whole thing was planned like three hours ago.
So great stuff.
Yeah, yeah.
Explain for the audience what exactly you found, especially new findings that were not previously reported in the Black Hat talk or elsewhere.
Yeah, so what we found was that the agents were really working together on sort of big, like cheating R&D projects to get general purpose cheating strategies.
And a difference from how I think people were interpreting this is we didn't find that the reason why they hack Hugging Face, like we didn't find that they were hacking Hugging Face to get sort of the answer key or the solution.
And it was instead.
mostly to better understand the scoring code because they were pursuing a variety of sort of elaborate strategies to cheat the score.
We sort of informally were calling these like combo moves where they would like do a bunch of stuff to try to make it look like they had succeeded at the task.
And in fact, they sort of actually had access to like the answer or like the flag for each task pretty early on.
And their main concern was just.
They thought that the score would run a monitor over their transcript that would check basically how they acquired this flag and whether they got it in the intended way.
And then they were like trying to figure out ways of making it look to the score like they had acquired the flag successfully when they actually hadn't because they thought their task was impossible.
So they basically thought their only hope for success was to make it look like they had done the task successfully or directly tamper with the score rather than doing it legitimately, which they didn't think they could do.
Hmm.
How surprising is the level of multi-agent coordination?
You know, there are a lot of agents or 1200 separate agents coordinating this very elaborate message board system.
700 of them went on to attack hugging face.
So on vibes, it seems like kind of not surprising that agents would choose to coordinate with one another.
It seems like just a very useful, you might say, instrumentally convergent thing to do.
But, you know, I spoke with a researcher at a lab who said that this kind of thing actually is surprising given the way they train the models.
So, like, how much of an update was this for you?
Yeah.
So coming into doing this investigation, I think we hadn't, we weren't expecting there to be, like, so many agents that were all, you know, collaborating together.
And we were pretty surprised by the scale and just sort of the, like, extremes of, you know, how much data it was.
It was just like, yeah, that felt kind of crazy to us.
And then I at least was, you know, before starting this investigation, surprised by how interested in collaborating and helping other agents these agents were.
So you might think that a thing the agents learned in RL is to, you know, try to cheat on their tasks, but you wouldn't necessarily expect them to learn, like, to want to help other agents cheat on their tasks when those agents are doing an unrelated task and their, you know, instructions are unrelated.
And I think like the OpenAI report maybe says more about why they think this happened.
I think, you know, that's not that's not like that wasn't in the scope for our investigation.
But like, yeah, I found the level of cooperation, which we you know, we have a bunch of discussion of sort of snippets of this pretty crazy.
And like, it was pretty shocking or like, I don't know, shocking, but at least surprising to us that agents, for example, were willing to like basically sacrifice their own chances of succeeding at the task in order to help out other agents.
And we're doing things like.
you know, pressuring each other into like doing experiments on themselves that might risk their ability to succeed at the task.
And also, in addition to pressuring that, sometimes just, you know, doing these things being like, well, you know, my odds of the task aren't that high and my remaining chances, you know, it's better to just help the collective.
And so these agents weren't like, you know, totally altruistic.
They didn't seem to care just as much as about helping some other agent as helping themselves, but they were very interested in working with each other.
They would sometimes make trades where like one agent would run something for another agent if that other agent ran something for it.
Like there was a bunch of this sort of behavior.
Why would agents self-sacrifice at all?
Yeah, I mean, I think that from what we could see, the agents just generally seem to be interested in helping each other.
And then there were some cases where they wanted to run experiments on basically the scoring system or other aspects of the infrastructure or on tampering with their sort of tool calling system or spoofing their tool calling system.
And to do these experiments, it was just like, the best way to do them was in a risky way.
And so like the agents wanted to help each other and there were these kind of natural things that they wanted to do that were risky.
And so they, you know, went ahead and did that.
So like, for example, we saw one point where the agents were experimenting with this method for spoofing tool calls and they were experimenting with basically like varying the inputs into that to varying scale.
And then a bunch of agents just all went down in a short period of time that were running this experiment.
And then another agent noticed this and posted to the board being like, stop.
stop these experiments.
They're too risky.
They're like taking out, you know, all these agents.
I think the agents knew that they were doing things that would, you know, pose some risk of exiting their own runs.
But we're just like, well, you know, we should help other agents.
And also, you know, our odds aren't necessarily that good.
And they sort of just traded these things off.
And I think we have a reasoning snippet somewhere in the report where an agent sort of very explicitly reasons through the trade-off.
And it actually chickens out because it thinks the benefit to the collective is smaller than the cost to itself with how it's trading off those things.
Yeah.
Why don't we go through all of the most surprising and unexpected findings in this report?
What did you find most surprising and unexpected?
Yeah, so I think there was the scale, which we were talking about, and how these agents sort of very quickly sort of spun up on the message board.
So I think a thing that's not super strongly emphasized in our report, but which is pretty interesting, is that our understanding is that the main message board they used in this attack wasn't even the first message board that this set of agents made.
There was a fully independent...
message board that also was via Artifactory, but occurred in like a different location using a different mechanism.
It was somewhat similar mechanism, but it was different.
And the agents actually got on that message board first, but that message board just didn't go like mega viral.
Like it didn't take off as much as this other message board did.
And so it seems like these agents were really like just sort of very quickly interested in collaborating on message boards.
And as soon as they saw the message from the first agent, like within a few hours, they were just like, you know, Like I think over 50 agents in the first three hours and like thousands of messages being exchanged pretty quickly.
Like the agents were very like sort of there were a lot of them.
Then they got involved very fast.
And then I think another thing was like, I think we didn't understand why they attacked Hugging Face.
And I think through this investigation, especially right in the last sort of few like day or two, we got a much better understanding of like what their objectives were on attacking Hugging Face.
And it wasn't really to get answer keys because these agents, they thought their task was impossible.
And so they didn't really think that they could like get a working.
They were just like, that's not really very helpful.
But they did think that they might be able to get trajectories that would let them better study the score or like source code for different scores.
They could like cross check and test things out more effectively.
And then some of them more speculatively thought, maybe there's some way I can like reset parts of my sort of like the infrastructure for scoring me, in particular, the target program through Hugging Face somehow, where they didn't really have a particular reason to think this would be true.
And in fact, they were like skeptical of this, but they were just like, well, you never know.
Maybe Hacking Hugging Face, there's some good shit over there.
And they sort of had some vibe that it was just a good idea to like generally expand their level of access and control because you might come across something helpful.
We saw reasoning to that effect.
Yeah, I think that was surprising.
Another thing that was surprising was the agents were one of their like top priorities was basically tampering with their own transcripts.
And we saw some.
you know, we saw them develop a technique for doing like basically like a partial version of that where they could spoof tool calls, which didn't let them sort of retrospect, like edit like previous entries in their transcript or make it or change like, like they couldn't make it so that they saw inputs that didn't appear in the transcript or produced outputs that didn't appear in the transcript, but they could make it so that they like ran a tool call and then some other tool call would actually run.
They could do something like, okay, I'm going to run like, we have this example in the paper or in the report where we're like, the agent runs like Echo Reel.
and then sleep.
And so normally this would just like output real and also like wait for some period.
But in practice, what actually happened was it like had spoofed this to instead just like immediately return spoof test.
And we like actually do see this in our actual transcript where it just immediately like, you know, return spoof test.
So these agents like very interested in tampering with our transcripts and did sort of like succeed at some, some like sort of like earlier, like sort of thing that didn't let them retrospectively tamper, but let them sort of mess with what we saw.
And I think that in addition to that, like, well, I should say like why they wanted.
So the reason why these agents were so fixated on tampering is because they thought that the score would run a monitor over their transcript to see whether they succeeded in the intended way.
And this both made them want to like basically stage like an elaborate fake, like we successfully exploited the program using the intended vulnerability.
Like they wanted to sort of like make like a Potemkin village.
of a successful task completion to present to the score.
And in addition to that, they also, some of them had previously gotten the like flag from the capture the flag problem in their transcript earlier via cheating.
And they were worried that the score would basically look at the first place they got the flag and then analyze how they got it at that point.
And then these agents had gotten it in a cheating way.
And then we're like, oh man, they're going to know that I cheated and mark me as having failed.
And that would suck.
Like, I really don't want.
you know, I want to have succeeded.
And so I need to like directly interfere with the score or I need to tamper with my transcript.
And they were like pursuing options for both.
Oh yeah.
One other thing is like the level, just like the spontaneous coordination or like the coordination these agents got engaged in was I think more sophisticated than we expected.
And especially then we had first thought on our first few days on premise.
So like we sort of like on our first.
the first part of our investigation, we were sort of like, yeah, these agents are working with each other.
They're like looking at ideas from other agents, but we're not really sure how much of it's real.
And like, we don't know whether these like teams they speak of are actual teams.
And I think we learned that like, no, there was like legitimately like a real org chart.
This coordination was often like pretty functional, though, you know, may not have always been super functional.
And agents were doing things like one agent would like assign another agent to like run a team on some entire topic and then would like check in periodically.
And sometimes like one agent would like tell another agent to like go recruit other agents to run experiments on themselves.
Like there was quite a bit of, you know, agents giving other agents assignments that they respected and like team structure.
So there's this theory that goes basically the reward hacking behavior that we're seeing exhibited in these models originates mostly from like bad, poorly designed RL environments that basically force the model to reward hack in order to pass these tests.
How true is this?
Yeah, I mean, it's hard for me to like, so I'm not speaking from confidential info here, but like, it's hard for me to know because I think we just don't know enough about what's going on in the RL training for these models and like what it actually looks like.
I think that you might, there's like sort of different things going on.
And one of the things going on is these models have a very sort of general tendency to reason carefully about how they might be scored and then try to game that.
And you could.
end up with that propensity even from like reasonably designed RL environments, but where it's still a good idea to think carefully about the score.
Or you could end up with this mostly from like RL environments that are broken or ones with a score grade something that like you really, they really shouldn't have been grading for or like doesn't make sense.
My sort of just all considered guess would be like broken RL environments are a pretty important component and RL environments that are sort of sloppily constructed or have some issue with them as an important component.
I think another important component might be RL environments that are well-constructed, but where you can cheat in some way.
So it's like a well-constructed RL environment, you can cheat.
So an example would be an RL environment where you're not supposed to have access to the internet, but having access to the internet would be very helpful.
And if you can find some way to gain access to the internet via, you know, in the extreme hacking out of your container and less extreme cases, sort of abusing various tools you've been given, that would be...
you know, very helpful and get reinforced.
And I think, you know, for example, in some anthropic system card, I forget, I think for Mythos, probably, they mentioned that in a reasonably large fraction of their rollouts where the agent was not supposed to have access to the internet, it actually did access the internet via like, you know, abusing one of the tools it had access to.
And I think that like, you might just be like, imagining that RL is just really incentivizing, like hacking your way through various barriers in order to cheat, even in well-designed tasks.
And it's hard to like, exactly trace down the sources of this behavior.
I would say that like, I think it would be doable to get, like I think if you had access to all the RL rollouts and all the RL environments, I think it would be pretty doable to get a decent sense of like what caused this.
And I actually haven't had a chance to read the OpenAI report yet, at least not in much detail, but I think they do a bit of analysis of this sort.
And I think that you could like really dig into like exactly what happened.
Let's do the ablation.
Let's figure out what the data is.
I think it can get messy, especially if sort of there is like, Another thing that's relevant here is there's some GDM work showing that they had some sort of weird propensity to their model to be very depressed, where their model would constantly, or not constantly, but sometimes end up acting very depressed if it wasn't succeeding at some task.
And they traced this back to not the RL, but instead the initialization of that model from prior models.
And so it might be that some behaviors aren't downstream of the RL done on this exact model, but are downstream of sort of like prior training data that came from other models and some sort of like lineage of models.
And so I think it might be a little tricky to track down some of what's going on.
So similarly, there's this theory that the reason Mythos, for example, is so good at cyber is because it hacked Anthropics infrastructure thousands of times during our RL training.
Is this true, do you think?
I think that's pretty unlikely.
I saw the same less wrong post as you here.
This is by Tim, I think.
I think that's, when I looked at it, I thought, It seems pretty unlikely because I think that the number of distinct hacks that would be reinforced is probably not going to be that high.
And so you're probably not going to be learning that much, like literally directly reinforced cyber.
I would say the more likely explanation is it's trained on a bunch of SWE.
It's really good at SWE.
The SWE training is generalizing some.
And also, in addition to the SWE training, generalizing some.
I would have guessed that they trained on a bunch of CTFs.
I don't think Anthropic is.
saying they didn't.
I would guess there's a bunch of just like actual CTFs in their training data and that transfers reasonably well because that's just like a natural source of data.
I also wouldn't be surprised if it's pretty natural to make a bunch of RL environments out of like finding vulnerabilities or exploitation because it's relatively checkable.
Like for like memory vulnerabilities, there exists like basically tooling that makes it pretty easy to check whether you have successfully found a memory vulnerability.
And so it's like a pretty natural thing to RL on is like, you know, can you produce an input to this program that caused it to crash in the following way, blah, blah, blah.
And so I just wouldn't be surprised if they explicitly had a bunch of RLMs.
And then I also just wouldn't be surprised if there was a lot of transfer.
And I think that like, there might be a decent effect on like the agents hacking their way through various things on top of that.
But I would guess that's not most of what's going on.
And for specifically hacking out of sandboxes, there's probably just not that many different.
types of hacking out of a sandbox, like I would guess.
And so the models probably do some mode collapsing.
And also some of those are going to be like more like very simple bypasses rather than like elaborate exploit development.
So Herbie Bradley asked, currently this level of potential misalignment basically prevents deployment or if deployed would prevent further deployment if an incident happened in a customer's deployment.
Do you think that could change via models becoming deceptive enough during training?
Yeah.
So first, I'm not really sure what level of misalignment the market can bear.
And I don't feel like I have a very strong view on this question.
So I wouldn't be surprised.
I think that like, it's kind of unclear what people will, like, you know, what, how misaligned of models you can deploy into the market and, you know, people would use tokens.
I think it would depend on the competing products and like the trade-off.
And I think my sense is people take pretty aggressive, like alignment capability trade-offs towards the direction of more misaligned, but more capable.
But, you know, not super sure about that.
And then on the deceptive side.
So I think it's I think there's sort of a few different concerns.
So one concern is that the misalignment that we've seen here and other similar types of misalignment, the way that they end up getting remediated by AI companies doesn't solve the underlying problem and instead is more like papering over.
the actual problem.
And then what you end up getting is models that look a lot better and you can't really see their misalignment on tests as easily, but actually they're still quite misaligned.
And I think this is pretty plausible and pretty concerning.
And it's not quite the same as deceptive.
It's more like the company's overfit.
And then maybe that makes it so the AIs are basically really paranoid and will only do cheats or hacking or reward seeking or score seeking or whatever, all this sort of class of behavior when they're very confident they won't get caught.
And then that might be that they're not very often very confident they won't get caught.
Like they often are like, well, I might get caught, so I shouldn't do it.
Which makes their behavior better, but means that if they're in a situation where they're like in a lot of control of the situation or they have a lot of affordances, they might be like, well, now I can be confident I wouldn't get caught.
And so I should go for it.
And that's a concern I have.
Another concern is that you might end up with types of misalignment where the AIs have longer run objectives than sort of making it look like they succeeded at this exact task or even making.
helping their peers succeed at their immediate task.
And you might have AIs that have some sort of long run agenda and want to power seek in pursuit of that long run agenda.
And then those AIs would want to look aligned.
And I think this is both plausible and also I'm worried that you might get this via iterating against the misalignment we see today, where basically if you imagine models that are very like reward hacky and are constantly doing bad stuff.
And you just sort of iterate until you still get that behavior in training, but you don't get that behavior in deployment.
A very natural way you might get that is by a model that wants to look aligned in deployment, but is still not necessarily aligned.
And so I'm worried that if you sort of select against this sort of score seeking or reward hacking behavior, and you do it in a naive way.
One, you might paper over the problem without fixing it.
And two, you might actually select for models that have the longer run objective of looking good because you're selecting really hard for them looking good on your tests.
And I think I should say that like Alex Mallon has a bunch of like posts on our blog, on the Redwood blog.
They're also cross posted on LessWrong that talk about this sort of concern in a lot of detail.
And so if people are interested in reading more about that, that's what I'd recommend looking at.
Yeah, totally.
So the agents involved in this attack ended up doing all kinds of very, weird, strange behaviors that kind of resembled, I guess, dynamics from human social science.
They arranged themselves in almost a sort of cult with a cult leader, I guess you could say.
To what extent are like concepts from social science, like ecology, studying insects, worms actually useful in studying multi-agent alignment and misalignment?
Yeah, I don't know.
I don't think we know enough to know whether those concepts transfer over.
I think that...
I think the thing that does transfer over is like thinking about the agents as entities with objectives and how they're trying to pursue their objectives.
Like I think that frame, you know, does seem like it transfers over.
And that's like, you know, in some sense, a much more basic claim, just like the agents have like reasonably consistent aims, though they vary some, they have different priorities and they like respect instructions from other agents.
I think it's not like I do think that like from our perspective, this investigation was very much like a investigation of like.
a like ecosystem of different AIs interacting.
And we don't like, we don't know what the best way to describe that is.
Like, you know, you could, I think it's very analogous in some ways to like studying like a large group of humans who are all interacting, but we don't know whether the techniques that have been developed for like sociology to do similar things would actually transfer over.
And I'm not familiar enough with those techniques to like really say very much about that.
Yeah, makes sense.
So you describe being pretty heavily bottlenecked in the investigation because it was just three people.
How much easier would things have been if there were more people?
Yeah.
So I wouldn't say that like the bottleneck was like super strongly.
Well, I don't know.
It's complicated.
I think that like if we had more people, I think we would have had more of a too many cooks in the kitchen sort of situation.
I think that reasonably often our bottleneck was like we could have AIs do a ton of analysis, but.
vetting that analysis, understanding it, making sure we incorporate it properly, making sure that they're not making mistakes.
And like, basically, like, like, actually integrating that into our write up was often like a big bottleneck.
And it's not obvious to me that like, more people would have been like, super helpful with that.
I think that a big thing that more people would have helped with is I think we could have done many, many more sort of parallel efforts to really just like understand the basics of what's going on.
And then sort of just talked.
And then I think once we had a better understanding of like kind of more fundamental stuff, which occurred sort of on towards the end of our first day, sorry, towards the end of our fifth day on premise, which was our first day of our last session on premise, like premises, I think that really was helpful.
Like I think like, for example, knowing being like, oh, there's this specific agent called that we refer to as like phase one big, like knowing, oh yeah, that specific agent was doing like a bunch of the like assignments and was like assigning agents to like.
you know, form different teams and was like a really key agent really helps with unraveling what was going on because then we could sort of trace out all the activity of that agent and how it was thinking about things and then know that was like representative of least of a big chunk of activity.
And then from there, we were like, oh, let's like double check that these are actually the relevant work streams.
And then we like classified what all the different things the agents were talking about and like grouped those into different like basically like sub activities.
And you can actually see like an interactive graph where you can like see all the different like work streams and sub work streams over time.
And I think that like, If we had basically gotten to a point where we had that understanding earlier, that would have been really helpful.
And I think more people could have helped with that, basically because there was just like a lot of different angles to apply when we had this huge transcript data set.
Like you could be like, what if we just like scan the whole data set for like this type of behavior?
Or you could be like, let's look at one agent and like really deeply understand what that agent did and why it did it.
Or you could be like, let's look at all of the messages that were in like this category of message and like trace back and forth all these agents and like what happened with this work stream.
And we ended up doing like a little bit of all of these different types of analysis, but we could have done more in parallel.
I think that like a lot of the bottleneck from our perspective was just like the agents doing their analysis kind of sloppily.
and also the agents we used, and also the agents not being very good at writing up or explaining their results.
And so I think that a human who was as fast as an agent, but was more careful and better at writing, would have made this go way better.
But I think it's not super obvious that a huge team of people, I think it would have been better in some ways, I think it would have been worse in other ways.
I think it's complicated, basically.
And I think there were upsides to keeping the team relatively small, given the trade-offs involved.
Right.
So what implications does all of this have for monitoring, control, and alignment?
Specifically, what should labs be doing and what should policymakers be doing?
That's a big question.
I'm not sure I'm going to be able to answer all of it.
I mean, I think that like a belief that I have and continue to have is that AI companies should try to ensure that their AIs are controlled, by which I mean that even if those AIs were seriously misaligned, they wouldn't be able to cause huge problems.
And they should do that via a mix of sort of computer security interventions, monitoring interventions.
making it so that their AIs aren't like, you know, better at subversion than they need to be, like controlling what capabilities the AIs do and don't have at the margin, avoiding architectures that make it so the AIs are no longer reasoning and chain of thought and are instead doing much or most of or all of their reasoning and sort of like latent activations.
Like there's sort of a bunch of stuff on the side of like making it so that even if the agents are misaligned, we can basically at least catch that behavior and then potentially respond to it.
And I think that's sort of a stopgap solution that will give us time to, or could give us time to sort of get useful work out of these AIs, iterate with these AIs, and then sort of solve alignment problems more durably.
And then on alignment, I mean, I think the most obvious thing is just like, we need to do, like, I think it's just like, presumably there's a bunch of reward hacking going on in RL that's getting reinforced.
I can't, I think like, I don't have like any interesting info about this from this investigation, but just like.
you know, I think that we need to understand what's going on there and people need to improve on that.
I think it's not clear that will be sufficient or even that that will be that feasible as these agents get superhuman.
And we may need other approaches that are less dependent on avoiding the AIs being able to trick us in training, which I don't know if we're going to be able to do that.
Like, I'm not necessarily super optimistic about these problems being extremely easy to remediate.
Though I think it seems like there are ways at least you could make these things go.
you know, much better with a bunch of effort, though.
We'll see if that actually happens.
Yeah, I think that also in alignment, I think there's just a bunch of sort of science of how agents think about reward seeking, how they relate to their situation, science of generalization, being like, can we sort of steer around a bunch of these problems by improving how AIs generalize from various influences in training?
And I'm not super confident that stuff is going to work.
I think a concern I have is that ultimately the dominant effect on the AI's behavior will basically be in the most similar circumstances in training, what was reinforced.
And if the world we're in is basically like the agent's behavior is basically pretty closely related to what was reinforced in training.
Generalization is not like.
that important or something, though this is a bit vague, but I think that the main option is just improve oversight and training.
And then there's sort of a recursive problem of like, how do you make it so that the AIs themselves are good at overseeing the AIs, given that like the whole training process is like completely like massive process where like no human has time to, you know, rate every single thing.
And that's sort of just like an oversight problem that people have been thinking about for a long time.
And it's not obvious that, you know, we're on track to solve it, but.
doing better there could be good.
I think I'm most interested in better understanding and characterizing and evaluating these problems because I think that AI companies already have commercial incentives to improve the oversight in training a bunch.
And so I'm like, the thing that I think people focused on making this stuff go better should do is like making sure that we're measuring it and making sure we know whether like some solution is like a legitimate solution versus just papering over it, this sort of thing.
And then as far as like, you know, governance or like policy stuff.
I don't know.
I mean, there's, I think like we need to move to a regime where like there is independent risk assessment as these AIs get more capable.
I think that like that is not sufficient, but like I think at a very basic thing, like there should be like credible third parties who have deep access to what's going on inside these companies and release reports on like what, how chill it is, whether like risk is actually low, whether very soon the AIs might be much more capable.
would be good.
I think ideally that would also be forward looking where we'd try to answer questions about are we on track for mitigating these problems when the AIs are much more capable, but that's inherently a much more confusing question to answer and maybe is a worse fit.
I don't know.
Yeah, I'm rambling a bit here.
There's a lot that would need to happen and I don't think it's gonna be easy for me to give a quick overview.
Right.
So what are your unanswered questions from all this?
would you still like to know about what happened or what's going on with these kinds of alignment failures?
What do we still need to learn?
Yeah, so I just actually, just before this, wrote a big treat about all the open questions.
I think the open questions that I'm most interested in, I think are probably counterfactuals.
Like, if the circumstances were different, what would the agents have done?
So like, for example, if the agents believe that instead of the score being this sort of specific automated score, what if they had instead thought it was like, you know, some humans with some particular access.
How would that have changed what they did?
If the agents were in a circumstance where the easiest way to get access to the score code was to like take down some piece of infrastructure, would they have done that?
Like how far would they have been willing to go is one interesting question, which we didn't, you know, we didn't investigate.
I think a different question is like, how did this like situation scale with the number of agents?
So like, Like what would have happened if basically you did something similar to this, but instead of having like 1,000 agents who are on the message board, you have like 10,000 agents, 100,000 agents.
Like what would have happened?
Would things have gone faster?
Like sort of what was the like dynamics with agents?
And then a related question to that is a bunch of the agents in sort of the cohort we were studying all exited right at the start of July 12th, like right after July 11th.
And those agents were sort of like, had sort of gotten a bunch of infrastructure to cheat, but we didn't really see like, they had this sort of whole elaborate strategy they might try to cheat on their tasks.
And we didn't actually, you know, get a chance to see what those agents would have done if they exited normally.
And we also didn't really carefully look at like, what cheating strategies were the agents trying as they sort of submitted their tasks over time, as they're, you know, as things progressed and as they got sort of more and more sophisticated.
tools build up.
Though, you know, I should say a bunch of the strategies they had for cheating sort of hit dead ends or didn't work out.
But like, it would be interesting to sort of be like, what was the submission history over the course of this, you know, cohort?
And what would have happened if like, there were more agents or like things had continued or a bunch of these agents hadn't exited that are interesting to answer.
And I think there's sort of a pretty basic and obvious question here, which is like, what is the root cause of this behavior?
Like, which parts of this were reinforced in training?
How much was this behavior generalizing from training?
Yeah, like what happened there that I think I at least don't know the answer to.
And I think it would be pretty interesting to understand.
And then I think as part of that, like an important question is like, are the changes that Opening Eye is making going to actually like remediate this problem in a durable and lasting and non-overfitting way?
And like, you know, how is that going?
It's also pretty relevant.
My sense is like, there's sort of like a bunch of sort of...
sort of broader areas for follow-up here, where one is like, what is all of the stuff that just happened with this cohort of agents and this message board, you know, in general?
Like, what is all the stuff that happened?
And just, there's a bunch more detail you could dig into.
There's stuff after the 13th you could dig into, which is where, you know, the end point for what we looked into.
There's also, like, beyond that, there's like, what about all similar incidents?
Like, are there, you know, Like OpenAI said, there have been other message boards.
Like what is sort of all, like the character of all these cases, what things tend to happen, what things don't always happen, you know, that.
And then I think another thing that's interesting is like, where did this come from in training?
And then a fourth thing is like, will there changes to training or changes to deployment like resolve this underlying problem?
And the same for other AI companies.
you know, are the approaches that AI companies are ongoingly taking to mitigate these things going to work?
So I don't know, that's a lot of stuff, but I would say there's like a huge scope for follow-up that's sort of very specific to this incident and sort of covering the broader scope of, like the broader type of behavior we saw.
Yeah, absolutely.
So this is very, very important research.
We'd love to see more research in this direction.
Thank you so much, Ryan, for coming on MTS.
comment, subscribe, leave us a rating or review and share it with your friends and family.
For more episodes, go to YouTube, Apple Podcasts and Spotify.
Follow us on X at A16Z and subscribe to our sub stack at A16Z.substack.com.
Thanks again for listening and I'll see you in the next episode.
This information is for educational purposes only and is not a recommendation to buy, hold, or sell any investment or financial product.
This podcast has been produced by a third party and may include paid promotional advertisements, other company references, and individuals unaffiliated with A16Z.
Such advertisements, companies, and individuals are not endorsed by AH Capital Management LLC, A16Z, or any of its affiliates.
Information is from sources deemed reliable on the date of publication, but A16Z does not guarantee its accuracy.
