# AI Plateau: Strategy, Cost, and Knowledge

**Podcast:** Dev Interrupted
**Published:** 2026-06-26

## Transcript

So, Andrew, what do you think?
Are you going to let mid-journey scan your body?
Well, maybe next time that I have to go get an MRI instead of having to lay down inside of a noisy claustrophobic machine, I definitely think it would be a little more, you know, like luxe spa experience to step into a shallow pool of water and then just like have a whole bunch of sound.
beamed at you.
I can't believe that's the direction AI is going.
I know.
Actually, when I saw this story, and for those out of the loop on this one, so Midjourney has announced a spa that they want to launch that has these new body scanners that they're inventing that use echolocation like a dolphin does and puts you in a vat of water and can scan you.
And when I first read it, I was like, it's not April Fool's, is it?
I was like, this feels like a joke, but yeah.
At the same time, I'm also like, but this is exactly where AI needs to be applied.
It's like, I've been saying for a while now that like the models, model advancements really aren't going to be the thing that changes the world anymore.
It's going to be like how we actually apply them in the real world.
So I don't know.
I mean, if they pull it off, like it seems pretty cool.
I think it's a fascinating application.
They're making some pretty good, bold claims.
They are making some bold claims about its accuracy, about its speed, also about its ease of use.
And that literally just works by stepping into a shallow pool of water.
And then a ring of sensors, they, like a dolphin, like you rightfully called out, they use echolocation to do imaging and then use that to create the end image.
So I guess in a way, it's almost generating an AI image or an output.
based upon the inputs of those echolocation receptions.
And so that's like a really fascinating flip of like AI on its head.
It also is a really huge step, I think, in like medical imaging.
We've actually, hilariously enough, talked about AI and medical imaging here on Dev Interrupted, I think about a year ago, because medical imaging is actually the realm of medicine that just really rapidly saw a lot of huge AI adoption.
And high success rates.
Because it turns out the AI is just way, way, way better at picking out nuances and MRI images and CAT scans and the like compared to even a really skilled physician.
So the idea of applying this to do a real-time scan of the body could just make medicine way more accessible for folks.
So it's a really exciting application of the technology.
Maybe a MidJourney Spa coming to your town sometime soon.
We'll see.
We'll see.
In the meantime, welcome to the Friday Deploy brought to you by Linear B.
I'm your host, Ben Lloyd Pearson.
And I'm your host, Andrew Ziegler.
And today we are covering, will the best models soon be out of reach?
Loop-driven development, the flawed search for a universal language and some wizard in their pond.
But we'll get to that one.
But first, I want to talk about AI models becoming out of reach in this new article from...
Friend of show, Steve Yege, titled The Flat Curve Society.
So, Andrew, what do we have here?
Okay, so this is the latest missive from Steve Yege, who you know we've been following nearly religiously here at DevInterrupted, about how agentic engineering has been transformed.
He's the man who brought us Gastown and started to introduce us all to the idea of autonomous coding factories at the top of the year.
So his opinion is definitely one to watch, and don't worry, we're watching it for you.
And here's the latest of what he's saying.
He's talking about the flat curve society, and this speaks to the idea of of what is the end result of once everyone gets into a Gastown way of working and we enter a Gastown type of connected world?
That's what, you know, Yege and his followers, the other folks building and using Gastown and Gast City have been building towards.
But along the way, we have to prepare for a huge, huge change in the users that are entering this space.
And this is what he calls the world of mediocre users and mediocre outputs and how to protect against them.
But more importantly, we also enter this realm where everything gets optimized.
so efficiently and domain expertise becomes truly the last island that you stand on, that even that starts to evaporate from underneath you.
He calls this the discernment horizon.
It's kind of a scary idea in which the idea is that you...
The machine and the process you've built is so efficient, and the work starts to climb so far above a realm of understanding that you're able to keep in your head, and then you lose your ability to efficiently grade the output.
So the ceiling on these systems, it turns out, is human-based, not the velocity of the gas town churning underneath.
And it's a lot closer to where we are now than I'll...
I think folks are giving it credit for, and that's what Yege calls out in this article.
And he also really calls out some of the practices and trends that we've been seeing in the industries, ones that we've been covering here pretty closely on Dev Entrupted as well, and at Linear B, stories around the build versus buy narrative.
And this is somebody who, this is the king of build it all himself, right?
And even he is joining the narrative and saying, Building it and then having a long-term success with the tool, using it at scale and with a large amount of users within a specialized org is a completely different basket of problems.
And that's why the build versus buy phenomenon has been so rapid and short-lived because as soon as it moves past that initial build, the entire idea falls apart.
So he calls out how we're in kind of this...
difficult to optimize new space.
If you were to think about like the curve of progression and adoption with Gastown and with Gastown-like development, we were in that huge rapid increase.
And now we're at that very top of the line where we're all optimizing for that zeroth of a ninth at the end of a percentage of efficiency.
The whole problem changes and it just becomes more fascinating to explore here.
Yeah, you know, most exponential curves stop at some point.
It's, you know, no system can continue exponential forever.
And, you know, we may be at the inflection point in this moment where that sort of rapid AI disruption is actually starting to exponentially slow down.
Now, it's not to say it's going to be slow anytime soon, but the rate of change is now beginning to slow.
And part of this is, you know, what's happening, you know, through all of this is like, You know, we have models like Fable that were taken away from public use because they were viewed as dangerous as weaponry.
And part of Yege's argument here is that one of the things that may contribute to this flattening is that we've gotten so used to these like exponentially better models that, you know, every month they just get substantially better.
But we may be at a new reality where those exponential increments don't get publicly released.
They actually are tightly controlled.
Most organizations may not even have access to the next generations of AI models.
You know, that has really big implications for all of us, really.
First of all, the models that, you know, us normal people out here have access to may actually be getting close to maxed out in terms of capabilities.
You know, if going further than this is too dangerous for public consumption, we may see limitations.
You know, we may be close to the maximum capabilities that we're going to see for some time.
And as you mentioned, Yege also thinks that, you know, the AI SaaS apocalypse has been delayed, you know, at least for now.
And we've been writing a lot about this.
We've been thinking a lot about this challenge recently.
And, you know, we even just published an article about how the AI SaaS apocalypse is a mirage, really.
And we experienced this firsthand by trying to use an agentic system to build.
the linear B platform and realizing just how quickly you spend a ton of tokens, not really actually solving the core problem that you have.
But then it also means that AI literacy is more important than ever.
You know, if we're not able to, you know, the pie in the sky view of models getting better forever and ever is that eventually we can just offload the vast majority of our thinking and our execution over to these models.
But the reality is that may not come to fruition.
And we may need to, you know, sort of compensate through our human expertise in the gaps that are in these widely available models.
But yeah, and there was also just some really great tips in this that referenced some training materials from Netflix that I really liked around measuring token consumption.
So, you know, we've ranted about how...
bad of an idea token maxing is and token leaderboards are, but it can still be a meaningful early sign of AI adoption.
In fact, we just ran a linear B workshop on this topic.
It's called Life Beyond Token Maxing, and we'll have a link in the show notes to it.
We really explored how token maxing is in token leaderboards.
You can start with them, but you have to quickly evolve to something that is more advanced.
But at least when you're starting out, you can think about it this way.
If you have one synchronous agent working for you nonstop, you're probably going to consume maybe 4 million tokens a day.
And then as you get into more asynchronous agents and multiple agents, your usage might scale up to like 12 or 15 million tokens in a day.
But beyond that, measurement really becomes a lot less meaningful.
You get a lot of diminishing returns because...
It becomes trivial at that point to just consume more tokens once you have these autonomous agents.
And that's really the point that you need to switch to outcomes.
You need to start thinking about what those systems are producing rather than just the fact that they are producing.
And that really, I think, is sort of at the core of this challenge of the flattening of the curve.
We can no longer rely on just throwing tokens at more expensive models.
We really have to be more conscious about being efficient with our token usage.
learning that when it's fine to spend tokens with reckless abandon, but then when you might need to use model selection to control costs or to optimize things.
And, you know, we're so used to this at this point to Yege's articles, just painting a picture of like complete chaos and uncertainty, like everything's being disrupted.
We're now operating in these gas cities and all of this.
I actually think this article is a really nice break from that.
Because it kind of closes with this idea that, you know, if we are plateauing, as Yege thinks, it's probably good for us in the long term, because it does give us a chance to get some breathing room, to start building for stability rather than just rapid iteration.
And, you know, maybe, maybe, just maybe, we're getting to a point where AI, the things we learn and the things we build stop becoming irrelevant within three months.
Like maybe we'll get more than three months out of the AI stuff.
we build and that would be pretty cool.
So yeah, it's a great article.
Again, from EIK, everyone should read this one.
All right, Angie, let's talk about the journey from test-driven to loop-driven development.
What's this article about?
Okay, so this is a fun companion for the Yege article because as you know, as folks who listen to this podcast most likely know, Yege gave us a really great language for describing the levels of agentic engineering and working with these tools, particularly with sub-agents at scale, where you can slowly watch the progression from things like, you know, we've seen this journey from like autocomplete to suggestion to auto-accept to you're not even looking at the code anymore, all the way up until you're managing.
running orchestrators at scale that are doing all of the subagentry underneath.
And so this article looks at how, what maybe the next step on the stepping stone there looks like, but more importantly, understands and acknowledges how each of those steps actually kind of encapsulate each other.
So all of those things I just described, they're really just zooming out one level further from the level before.
It's not that the level before doesn't exist anymore.
It's just that where your attention is going is different.
And so the user experience, the coding experience that you curate for yourself ultimately just takes a different shape, right?
And this one argues that you want to arrive at a loop-driven development pattern.
That's what sits one level above kind of what we would...
Label was the traditional top of like the Yage chart of having your orchestrator.
And the loop-driven development is then confounded by time.
And this is honestly one of the most elusive parts about working with these agents at scale and on a regular schedule and a regular cadence is, okay, you've got your great harness.
You have your awesome process.
You can turn anything out from a spec, from an idea into an application really quickly.
And you can use your AI and your skills like...
tools to get recurring tasks and things that are constantly on your plate knocked out really easily.
Now, the next challenge is identifying on a weekly or monthly, these different loops that you live within your work basis.
What are the types of thinking?
that I need to be happening around me?
What are the assets and the end results that I need delivered to me?
How do I expect my domain expertise as somebody learning XYZ or working in this space to grow over that period of time?
And how are my agents and my outputs going to keep up?
This loop of learning, this loop of development, it happens on several layers.
It happens on a weekly basis, even a daily basis, all the way up to a quarterly.
yearly.
And I think that, you know, the argument of this is about understanding that that becomes your next challenge.
If you're an engineer who's listening to this and you feel that you identify with a lot of those traits that I was mentioning before of having the stuff at your fingertips and being able to use it, I challenge you to think about how much of that happens without you having to take a direct action to invoke it.
And what are the rules by which they come to your presence?
And how are you able to review them?
I think everyone would benefit from just imagining just for a day.
Like if you were to hire an executive assistant that was going to help you do all sorts of stuff, what were the things that you'd want them to bring to your desk for you to review?
And these are going to be from other experts within your company, right?
Like you want your legal guy to review all your legal stuff.
You want your VP to be giving you your engineering report, but you also do want it all come together in this one collected space.
That becomes the challenge with loop engineering is identifying what that looks like for you.
Yeah.
And if you're someone who feels like you might be behind the curve on AI literacy or you're struggling to get from where you are today to the next step on the sort of like AI maturity path, I think this article is a really, really good read for you.
And as you mentioned, we've been covering, we've covered Yege's model for AI maturity.
And I think this article is really in like...
A very, very similar thread, but what I think it really excels at is breaking down those additive skills that you mentioned that you need to progress through to level up from low AI fluency to fully agentic development.
Particularly, I would say that most people today have probably advanced beyond using autocomplete and even prompt engineering.
That kind of feels like old news to me at this point.
And I feel like the general zeitgeist right now is context engineering.
You know, that's really been like what we've all been thinking about for the last year or so.
And if that feels like you, this article has a really great description of how to get into harnesses, into feedback loops, because those really are sort of the next critical steps to getting to a fully agentic coding place for a while.
And, you know, and we've been saying this, we're kind of a broken record on this at this point that, you know, the harness.
It matters just as much, if not more than the models that you're using at this point.
And it really is the only way that you can reach a state where you can build those back pressure loops to fix all the issues that an agentic system may encounter.
So, you know, if you feel like you're stuck somewhere in the middle on AI fluency, this is definitely an article that you should read.
All right, Andrew, walk us through the failed universal language and how it explains why we're all struggling with AI output.
I love this article because it is one of those articles that reframes how you think about using these tools, but also just kind of gives a good mental reset on how you should be thinking about your own work and ways that you can improve your output.
So this is looking at our, I guess you could say, flawed search for a universal language of AI.
And really what this article is diving into is how we as engineers, as general knowledge workers, have traditionally up until now kind of used models in a comparative sense to kind of figure out what's the role with.
Everyone has developed an opinion at this point about what model they turn to for what.
And certain models maybe are earning reputations for delivering on certain types of work over others.
So you get these emerging quirks.
And we've all had our discussions and personal opinions about how even just the writing style of these models, these mainstream models, models that a lot of folks are using just differ slightly.
You get these really cool hands-on experiments like we've covered here on the show before, like AI Village, that's taking things like Gemini and ChatGPT and Anthropics Cloud and putting them all into a simulated chat together and doing goal resolving together.
And you can even understand how they are separately thinking and acting from each other.
And all of this nuance emerges.
What's underneath supposed to be just like the same kind of technology.
actually has a fair distinct amount of emerging traits within it.
And this isn't to anthropomorphize it at all.
It's still ultimately, of course, just a text prediction machine.
But the real...
gap or opportunity that this article calls out is that we use that to put them side by side like it's a race like it's a whole bunch of like cars about to run a relay and you're going to give them all the same prompt they're all starting the same exact you know whatever starting point and then you're going to see who's going to deliver quote the best output of based upon what you gave it and you're going to measure it.
And this is how folks have been seeing like, oh, which should I use?
Which model should I use to build this spec or which model should I use to actually implement this in production?
And people kind of choose their lanes.
Now, what you're throwing away in this world is actually understanding how those framings of those different models, what they call out about your original input.
And this becomes a useful signal because this.
Disagreement on framing often means that your own thinking is still unresolved, that your own prompt is ambiguous, or you haven't really committed to a certain direction yet.
So if you get a lot of deviation, that's enough for you to pick a winner.
What this actually is calling out is that there's enough missing information and bias, perhaps, within your initial starting prompt that is perhaps something that you're not seeing.
The opportunity becomes then, what was missing in my prompt and my instructions that caused this divergence between these three or four or whatever very competent models because that speaks more to your idea and how fully formed it is than the prompts they produce.
You and I, Andrew, we experiment a lot with AI models over time.
And I know particularly you, you've been, you know, the article mentions Quinn as one of the go-to models for this author.
And I know that you've been playing with that a lot in addition to other models.
But...
You know, I'm definitely guilty of like falling back to a daily driver mode.
You know, currently I choose Opus 4.8 for practically everything because I like it a lot.
But it's always been in the back of my mind that like I'm probably not picking the best model for every task that I'm engaging with.
I'm just going with what's convenient at the time.
And yeah, like sometimes, you know, I'll pick a less capable model just to reduce usage costs.
Like if I know that a cheaper model can solve it.
I may, you know, tune it down a little bit just to save some of my session window.
But I've never actually really given a lot of deep thought into why models behave differently based on the type of prompt that you give them.
And that's what this article really dives into pretty well.
And a lot of it just comes down to the linguistic approach that they take to it.
So, you know, these are stochastic systems.
They just predict words that need to come next.
And if they have a bias towards a certain type of linguistic approach.
It'll actually influence their behavior at their core.
But, and this can be a thing that, you know, creates, you know, it helps you identify uncertainty within your own thoughts, as you mentioned.
And I also think it's a thing that you can learn to benefit from, you know, because the linguistic approach of Quen may be better for like steel manning your ideas, but the polished linguistic approach of something like Claude might be better for delivering higher quality artifacts.
So, yeah, I think the advice here is really just to test the assumptions of the models that you're working with and use that to your advantage.
You know, if they have different opinions, it's an opportunity to use those differences and opinions to improve whatever you're working on.
All right.
Now I want to talk about building a local knowledge base in Google's open knowledge format.
And first of all, I want to point out, I had no idea that this was a thing.
Google's open knowledge format.
I've been doing it for a few, like a few months now, but I didn't know that's what we were calling it.
But anyways, welcome to the party, Ben.
Yeah.
Anyways, this is a lightweight spec for organizing knowledge as a directory of interlinked markdown files, which is, you know, it's a pattern that was really sort of like popularized by Andre Karpathy.
He's the one that I learned about this technique from and who I adopted it from.
And he's been advocating it for a while.
But the idea is you give your agents the ability to read, write and navigate structured knowledge without any sort of database or special tooling.
And this article that we'll link to, you know, it does a really good job at showing getting hands on with a lot of different tools.
You know, the way this works is you build a raw data set and you ingest that into your AI and have it build its own wiki with cross links and cross references to all the different resources that it needs.
What it really comes down to is compiled knowledge and plain markdown is actually a really durable and agent friendly way.
to give context to your AI systems, even better often than like vector databases, particularly when you're working with like internal documentation or team wikis.
And it's really just because Markdown is portable, you can version control it, and it's really easy to consume by both AI and humans.
So, you know, I wanted to include this article because, you know, I don't have a whole lot of opinions about it other than to say it's just a great, you know, yet another person out there that's operating in a very similar manner to how I've.
been operating.
And, you know, I just love seeing different perspectives on this approach.
So what did you think, Andrew?
I think it's just funny that the idea of folks trying to give or take or label credit for what's ultimately just like front matter on the markdown files, which is a practice we've all been loving.
And also, by the way, I want to call out that this arrives just in time for all of us to get wandering eyes for HTML as the new end place where all of this kind of stuff should be going.
Because HTML is actually a much richer way to express this information because HTML5 components are semantic compared to markdown, which lacks semantics.
not to mention that it's renderable in places for the human and the agent just as easily as Markdown.
But, you know, we're not ready for that conversation yet.
I think that I'm ready.
I just think the technology isn't ready.
We'll catch up.
There will be another, there will be an, there will be like an open, there'll be an OKF for HTML.
It'll be like OKF HTML or something.
And we'll be talking about this acronym like a year from now.
I'll probably have a tutorial out about it.
So just stay tuned.
I think that obviously, organizing your knowledge, putting it down into a durable store.
This is like baseline investment for anybody who's working as a knowledge worker.
And that's anybody on like a management side who has meetings with folks, which is most people, a lot of folks, especially that like work in our industry.
But also too, if you're delivering code, if you own a domain and expertise, it's helpful to write that stuff down.
Even if like you're writing it down somewhere that belongs entirely to you, like.
go you.
That's exactly what you should do.
It's the domain expertise locked up inside of your head.
It's very specific to you.
And for you to actually give it the depth and the color and the volume that you need, it becomes a little personal on some levels.
And so, you know, use this as a way to refine your outputs to the world.
But don't think that you have to play all your cards out there.
It's a practice that's important to pick up.
All right.
I wanted to end this today's episode on a really fun one.
And our producer, Adam, really hit it out of the park on this one.
And this article is titled The Wizard with the Very Defensible Pond.
It's an allegorical essay by Scott Werner where he uses a wizard and a pond and this traveling sorcerer and goblins and an apprentice.
Really, it's just an allegory for the disruption that, you know, a lot of established companies are feeling because of AI, I think.
And it's just it was I wanted to include it.
I don't want to say a whole lot about it because anything that I say would be a complete disservice to the absolute incredible writing of this article.
There are multiple moments where I found myself laughing out loud at it.
But, you know, and I think, you know, the moral of the story is that there are limits with what AI can and can't replace.
And we really need to be aware of where those limitations are.
And this allegory is just a wonderful illustration of it.
So what'd you think, Andrew?
Scott Werner, you are a creative genius.
This story is so fun and so relatable, and I really couldn't put it down.
Any allegorical story that manages to tell a lesson that totally sticks is totally obvious, and it's also dressed up as something we all know.
It's this fun kind of children's book kind of story with illustrations that go along.
Honestly, once you start reading this, you won't be able to put it down, and then you'll have lots of opinions, and how it applies to the world we find ourselves in will be immediately obvious to you.
So to Ben's credit, or to Ben's point, you know, I'm not even going to try to summarize it.
Just go read this really fun allegory, and be sure to share.
it because honestly, we need more kinds of essays like this.
This is a really short and sweet illustration of the pain and the reality that we're all going through right now.
And also too, just it has like some lessons built in as well.
All right.
Well, Andrew and I, we just wrapped up this really great session on token maxing and really just getting to a life beyond token maxing.
And we had a lot of fun running through it.
We got a lot of feedback, a lot of questions from the audience.
And it's very clear that this topic is really hitting a nerve right now.
And I think it really has to do with the fact that executive conversations all over the place right now around AI has really shifted from this, let's get everyone using it to now everyone's wondering, how much are we spending?
And is it actually worth what we're spending?
So if you missed the live stream, we have the full replay on demand over at linearbeat.io.
And, you know, the reality is that your CFO, they aren't looking at adoption rates or token counts anymore.
They want to see what all of that generated code is actually delivering for the business.
So in this session, we map out exactly where AI is shifting bottlenecks in your pipeline and how the LinearBee's Apex framework helps you measure what is really valuable to your business.
So we'll share a link in the show notes, but you can also head over to LinearBee.io to check out the full session.
Thanks for sticking around all the way till the end of the Friday's a ploy.
We always love sharing these news articles with you.
If you love what you heard today, the best way you can help us out is just to help us spread the word.
Share the video with your friends, the podcast, with whomever you think might want to listen to it.
It really helps us grow the show and we really do appreciate you sticking around and helping us make things better over here.
Find us out on LinkedIn, on Substack.
We'll see you next week.
See you next time.
