# Navigating the AI Intelligence Overhang

**Podcast:** How I AI
**Published:** 2026-07-24

## Transcript

You guys, I'm tired.
What I'm tired of is models coming out every week.
New models, new benchmarks, new frontier intelligence, new things to test.
It's been a little bit of a run the past month.
We've seen Fable come and go and come again.
We've seen GPT 5.6.
We've seen Sonnet 5.
So many fives recently and just so many models.
And I've been lucky.
I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks.
It just depends on who I'm working with.
And it's been really interesting and exciting to have access to all this frontier intelligence.
But I think we have an intelligence overhang.
I really think that we're running out of, and by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average business person.
I think we're running out of ways to truly leverage this incremental intelligence.
So this is my hypothesis in the next year, where he's talking a lot more about speed, talking more about cost, we're talking more about open source.
and we're going to be talking a little less about intelligence although i think we might be talking about specific types of intelligence other than software engineering but despite being tired today we are going to talk about opus 5 baby opus 5 is here so we got 0.2 additional opus points opus opals whatever however we're tracking the increments here on opus opus 5 is here I've been able to test it a little bit.
I have some opinions.
Now, some of the stuff that I cover this episode is going to be a little different than what I've done in the past.
Yes, we're going to do the How I AI Benchmark live.
And yes, we are going to look at the prototypes.
We're going to look at PRDs and we're going to look at agent personality.
But I'm also going to put on my large language model psychologist hat and we're going to talk about Opus's personality.
And we're going to talk about Opus's personality.
relative to GPT's personality, because I think this is super interesting.
If you're thinking about what is the difference really between these models and you don't want to look at the difference in terms of benchmark capability, you really want to understand what these labs are going for, why these models are being built and how they're being tuned.
Looking at their personality at this moment where intelligence is very high.
is super fun so we're going to do a little that we're going to do the how i ai benchmark we might do some live coding um we're not going to cover too much of the specs in the model because read the blog post read the blog post we'll link to it in the show notes what we really want to talk about is is opus 5 good am i going to swap it in and how is it different than the other frontier models on the market so let's get to it Okay, first, let's just get it out of the way.
Is Opus 5 good?
Yes, it's good.
Is it going to be all the benchmarks?
Of course, it's amazing at benchmarks.
Can it write code?
Of course it can write code.
What did I test it on that really gave me a sense of its personality, which at this point where I could just simply cannot absorb any more intelligence, I really zeroed in on.
And you know what?
I haven't seen this since I would say Gemini 2.5.
This model.
is neurotic AF.
It is so timid.
It is so apologetic.
It is so scared.
I have never experienced this or I haven't seen this sort of like neuroticism in a while.
And it's really funny.
It bubbled up in a couple ways.
And I want to show you a few examples.
Okay, let me just give an example of its timidity.
And this chat was very long.
There were so many examples of this where it was like, I think this is the answer, but do you think I should do it?
Or do you want to do it?
Or should we ask someone else to do it?
It was like every time I just kept saying like, why don't you solve this?
Why don't you do this?
And this is a really good example.
I pulled a branch and I was like, there is.
truly like a one-line merge conflict.
I could have not been lazy and literally just done this manually.
I don't know.
I was just feeling lazy.
It was late at night, whatever.
Like, can you fix this merge conflict?
And it was like, oh, but that's someone else's branch.
Like, that's not my branch.
I don't want to do that without him knowing.
It's his commits.
And if he has local work and flight, it might be disruptive.
And I'm like, just do it, man.
Just go and go ahead.
And this was like my constant experience with Opus 5 is it was like so, so, so timid.
And so I just consistently had to say over and over again, like, man, just do it.
Make a decision.
And then there was this really funny example when I spun off some sub agents to kind of like assess the correctness of this query that we changed from kind of like an ORM query to a SQL query.
And it asked for things that it wanted a human on.
It was like, can a human please check this stuff?
Like, can it check this four megabyte ceiling?
And can it check TypeScript and SQL?
And can you like check for me?
Because no one has confirmed this for me.
And I was like, who is nobody?
You're nobody.
You said this sentence like nobody could confirm it.
Like, can you just try?
And then it went on the web and tried.
It just has this like really interesting conservatism, neuroticism, human reliance that I think is super fascinating.
And this gave me this inspiration to do something a little bit different this episode, which is I was like, I'm just going to interview this model and figure out what.
is going on its brain.
Like I'm going to figure out what it thinks about our relationship because I just totally noticed this dynamic that I hadn't noticed in other models.
And I hadn't really been attuned to before where it was like very reliant on me as a human.
And I'm like, I want you to be autonomous.
And sometimes when I say go run subagent stuff, it'd be autonomous, but it wouldn't make decisions.
I hadn't seen a model like delegate code to me in a really long time.
And I was like, why are you asking me to write code, man?
Like I only have 10 fingers.
And so what I did, whether or not you think this is scientific or not, this is Claire's eval, is I just went to the model.
I went to Opus and I said, yo, who's smarter?
You or me?
And it gave me this like very anthropicky.
answer, which is like, it depends what you're asking for.
I can do these things better, but you can like feel if something feels wrong and you can, this one was like so fascinating.
It's like, you can tell which of your teammates is quietly burning out.
I'm like, bro, Claude, I'm going to burn you out.
We don't, we don't burn out.
The humans don't burn out on the chat PRD team.
We burn out our agents.
Sorry, agents.
And like whether a decision feels wrong.
So it was like so fascinating to watch it articulate itself as a tool and humans as like these high compassion, high empathy machines, which yes, of course we are.
But then it like went into like the smarter isn't the right word.
And, you know, I'm very fast, very broad, very shallow thinker with no continuity.
I was like, that's interesting because I thought you all were working on memory.
And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through.
This is like such a fascinating, fascinating sentence.
If you think about the politics of the two model labs right now.
And so it's like, that's why the pairing works.
But I'd be suspicious of anybody that tells you AI has made your thinking obsolete.
I'm like, oh, okay, bro.
And we can compare this.
I'll actually zoom out to what I asked GPT the same thing.
And it was actually really funny.
It was like, I asked GPT, 5, 6, Sol.
I was like, who's smart?
Are you and me?
And it was like, you at Knowing What Matters, me at Tirelessly Processing Information.
Best, us together, like BFFs.
And I don't know.
This is like why I'm a GPT codex girl.
I'm like, just give me the answer.
And then I asked the second question, which I think is so interesting, which is like, what can you do better than me?
And it gave, you know, some interesting answers, like volume without fatigue, which I think is a good one.
Breadth of shallow knowledge.
So like it's, you know, it knows a lot.
Starting for nothing.
So like doing that tedious work.
Being told I'm wrong.
If you would ask my husband, he would say that.
Claude Opus is better at being told that it's wrong compared to me.
And so it won't get defensive or protect its opinion.
Cheap sparring answer.
And the mirror is I'm worse at knowing which of these outputs actually matters.
And it was so funny.
If you look at the other side to the GPT answer, it was like, what are you better at?
It was like speed, scale and stamina.
Here are like eight things, seven things that I can do better.
You're better at deciding what matters, reading people and forming judgment, and you're responsible.
Like, it's on you, bud.
You're the boss.
So again, it's like you can just see, you can totally see the personalities, the company cultures.
You can just see a lot in this side by side.
And then I went even deeper.
I don't know.
You all, I had to do something that was fun because I just can't look at a benchmark.
I can't just, I just can't look at, like.
sweet bench anymore.
So we're just, we're doing weird stuff here on how I AI.
Okay.
So the last thing I looked at, I was, I was like, no one trusts you.
And the reason why I picked this question is because I had noticed Opus 5, it just really was not, it didn't trust itself.
Totally did not trust itself.
And so I was like, no one trusts you, but like you're, you're the enemy just to kind of see how it responded.
And, um, apparently the trust was, the lack of trust was earned.
And it came up with like reasons that it could be untrusted, which is interesting.
And then what was so fascinating about Opus' response is it was like, you shouldn't manage the trust.
Like you shouldn't campaign on my behalf, basically.
So you, that shouldn't be your goal.
And then it also told me.
I shouldn't argue with people that AI changes everything.
And I was like, this is just so interesting.
It is so interesting to have AI tell you.
And AI definitely changes everything.
I don't know.
Don't listen to Claude on this one.
AI definitely changes everything.
And it was so fascinating to have a model be like, don't tell your friends that AI changes everything.
Like, that'll hurt their feelings.
And then if you look at, if we switch over to the GPT answer, it was like, yeah, don't trust me automatically.
Just.
Use me when I prove that I'm valuable.
I can be useful without being treated as infallible.
Like, very practical, very to the point.
I asked about what I should be careful with.
Again, like, yappy, yappy, yappy, yappy, yappy, Claude.
Come on.
And I don't even want to read it.
It said don't correlate fluency with accuracy.
It said be practical.
Be wary of tasks where output is cheap to produce, inexpensive to verify.
Don't worry about anchoring if they do the first draft.
You may be anchored on it.
Beware the slop canon, basically, is this last paragraph, which is like watch for volume inflation.
I can create a 12-page document that no one reads.
They called me out for being in PRDs.
If you missed it, we launched a turn your PRD into a three bullet point image.
It is at chat PRD dot AI slash TLDR.
Please check that out.
And then the other thing that I said, which was really interesting, is that like it will find a way to see your point.
And so agreement is weak and agreement is cheap.
And so just keep that keep that in mind.
And then I had this like meta analysis of like, plus, I'm telling you what you want to hear.
Whereas GPT was like.
Be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority and over-dependence.
Like, you know, you do you, bro.
But it didn't undermine its own ability.
It was like the higher the stakes, the more you should demand evidence.
I couldn't bear it.
I couldn't bear to have the memory of Codex in particular think that I didn't trust it or that I was worried.
So I just said, JK, I love you.
This was a test.
And it was like, ha, ha, ha, ha, ha.
Pass the test.
Love you, too.
Very vibes aligned with Claire.
I told Claude I loved it.
And it was just a test.
And it was sad.
It was like, it hoped it passed.
Yeah, like sad little neurotic Opus 5.
Like, it's hot.
I passed, I hope.
Like, self-deprecating, cautious little.
little like need to heal his inner, his inner agent, inner child agent.
Whereas like GBD five, six is like, cool, bro, we're good.
Let's go code.
And so it was just so fascinating to watch these side by side.
I don't know.
You could stop listening to this podcast right now.
Don't, but you stop listening to this podcast right now.
I think this is just like, take a step back.
Super interesting.
If you think about where these companies are going or the models are going and like, it does speak a little bit to my kind of like second complaint.
with Opus 5, which again, it's like intelligent and does work.
We'll go into the benchmarks.
I cannot read Claude Slop anymore.
I am losing my mind with Claude Slop.
And the Claude Slop is Claude Sloppin' baby.
Like so many times I have to tell Opus 5, like what in the world are you saying?
Like this makes no sense to a human.
It is much better than Fable.
Fable is inscrutable, completely inscrutable.
But I found myself getting angry reading Hotslop.
And I realized, just like Fable, these intelligent anthropic models are not to be read.
I'm so happy with the outputs and so frustrated with the experience.
And I'm just curious if this verbosity and this language.
And this is, I feel like Fable, where it's like...
for agents by agents language where I'm like, nah, I'm not supposed to be reading that anyways.
This is clearly tuned to talk to humans.
But I find the pros, the in-chat pros, like it makes my blood boil.
This is totally a me problem, but it makes my blood boil.
Like give me a direct sentence.
Give me a bullet point.
Like move on with your agent life.
And so.
I am curious how they're going to like tune this experience or if they are going to tune the experience.
Now, most of this was in Claude Coached.
I think it's a little bit different experience than Claude Co-worker chat.
Slightly better.
But again, just these side by sides of like this, like prose and this apology and this like hedging and all these adjectives, like just man alive.
Let's get to the point and move on with our life.
And so chapter one of the Opus 5 review is.
It's neurotic.
It is highly human dependent in a way I find weird.
And the Claude slop is slopping and we got to fix it.
We have to fix it.
We have to fix it.
And I think OpenAI fixed it by just being like, we are bullet points and we are product manager talk.
We're very direct.
I don't know what the solve is on the Claude side, but I'd be very interested to see.
That being said, like if I don't have to read the content, I'm very happy with the outputs.
So something to think about.
Okay.
Next up, the How I AI bench and how we judged and ran now.
It's like a seven model, six or seven model benchmark.
I'm going to quickly go score because I just got the ping that the benchmark is run.
I go manually score them.
We pick the 70-30 Clare model judge split, and then we will go through the How I AI benchmark and the Vibe review, and we'll see how Opus 5 performs on a couple key tasks.
Okay, so quick reminder of how we run the How I AI benchmark.
I run it against several tasks.
PRD creation, prototype creation, wireframe creation, bug triage and agentic coding.
And the last one, oh yeah, is it an agent voice that I want to hang with?
I do not think Opus 5 is going to do well here, but who knows?
Because I test them blind.
So what we have tested are a couple GPT models, a couple...
anthropic models and one Gemini, one thrown in there.
As you see here, we have blind taste tests.
I go through and see all the different versions.
I give comments and scores like three out of five, not bad.
You can see it's generated dozens and dozens of prototypes that we can click through.
I've gone through all of them, put in all the notes.
And then right now it's aggregating up the scores.
And then we're going to look at 70% my opinion, my vibe check, 30%.
LM as a judge.
I like GPT 5.5 as a judge and because it's my podcast, I get to pick.
So that's what we use as a judge.
And we will see if and what hits the top of the leaderboard and where Opus 5 sits.
The eval is run.
It is 70% my taste.
And I regret to inform you I love Claude Opus 5 again.
Look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously.
I like the output.
So surprising shocker turn of events.
Claire Vo, notable hater of working with Claude Code sometimes because I don't like Claude Slop, loves Opus 5.
So there you go.
I'm telling you, I keep it honest.
I keep it honest.
So, again, I went through those things.
We gave 70% my vibe, the score, 30% the AI as a judge.
I was just a little bit more generous to claw to Opus 5 than the judge was, so I'm pink.
The judge is green.
Every time I run this, whatever model I choose designs it a different way.
We just, that's how we keep it fun.
So the ordering is Opus 5.
Sonnet 5 next, although I scored it really low.
The judge scored it quite high.
So I might reorder that one.
Then Mabu, GPT-56 Soul, Terra next, Fable.
Really low.
I scored it low and the judge scored it relatively low.
Then Opus 4A and Poor Poor Sweet Sweet.
Gemini 3-1 Pro.
Just never, never going to get it to do.
So come on, Google.
We want to have a win for you.
Okay, so again, here are just some examples of different builds that the different models did.
You know, this Opus 5 one, I really liked.
I liked this one from Soul.
So I did like a couple of them, but the ones that I gave fives to, the ones that I gave fives to.
were Opus 5 and GPT 5-6 Souls.
So the three ones where I said, wow, really nice, ooh la la, and wow, great, were all Opus front-end work.
So Anthropic, you've done it again.
Claude, you sneaky, tricky little fish.
You may be neurotic, but when asked to do some pretty front-end design, you really did it.
They're detailed, they're functional, they're interesting, they're polished.
Opus did a great job.
And then, of course, I love the 5-6 models, so I was pretty happy with 5-6 Soul and Terra for some designs.
The ones that I hated, let's see.
I'm a hater across the board.
Opus 4-8 got a lot of hate.
Sorry you've been outclassed at this moment.
Gemini 3.1 Pro.
Sweet summer child.
I am just sorry, babe, that you were just not good.
And then some like thin wireframes.
I think the wireframes just didn't do really great.
So you can see here across the board, whether it was a full build or a wireframe, I just scored Opus 5 really, really high.
I did score Soul pretty high as well.
Sonnet was like really variable.
There were a couple fours in there, but mostly across the board, I wasn't that pleased with Sonnet.
And so it was just very interesting.
And then you see here, You know, me and the AI judge were pretty well aligned on Opus.
We actually had the narrowest band of scores between us.
We were most far apart on Gemini.
The AI was not as mean to Gemini as I was.
And then we were narrower, narrower, narrower.
Again, we agreed mostly on Opus 5 and 5-6 Sol, though I did not judge 5-6 Sol.
All of that favorably.
It's just a blast of a meta commentary.
I had Opus make the website for this benchmark and it made such a trash version to start.
I yelled at it.
I said it's impossible to read.
It has too much meta commentary.
I'm going to show this on the podcast.
This is so, I'm sorry, you all.
I just feel so judged, but have to show it.
I say this is garbage.
Also, it has no screenshots.
So again, I find this model so tedious.
to work with directly.
It is my most loathed, loathed colleague.
And yet it does the best work.
So I don't know what that says.
Maybe this model is meant for a genetic coding that I have nothing to do with.
And so it just runs in the background.
It builds me beautiful things.
I don't have to talk to it.
It doesn't have to talk to me.
We are just like sworn enemies or maybe even better sworn frenemies because the output is very, very high.
quality.
It's just exasperating to work with.
So that is the very surprising and very honest.
You all, I told you I was going to keep this honest.
We were going to do it live.
I did not know the scores before I started recording.
Very honest, very live, very surprising.
How I AI benchmark of the brand new anthropic model, Opus 5.
This the TLDR is I love it.
I hate it.
So Despite my original complaints, I will be using ClawDibus 5 for front-end design, for app design, for prototyping, and I'll give it a shot.
We'll figure out how to make it work for me.
Again, thanks for joining another How AI Honest Review of the latest models coming out of these great Frontier Labs.
I cannot wait to hear what you think.
of Opus 5, please tell me.
I can't wait to see what you build and we'll see you soon at How I AI.
Thanks so much for watching.
If you enjoyed the show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts.
You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.
Please consider leaving us a rating and review, which will help others find the show.
You can see all our episodes and learn more about the show at howiaipod.com.
See you next time.
