# AI Judgment Models Reshape Enterprise Automation

**Podcast:** The AI Daily Brief (Formerly The AI Breakdown): Artificial Intelligence News and Analysis
**Published:** 2026-09-16

## Transcript

It's not every day that we get a new model to play around with, and it's certainly not every day that we get an entirely new approach to model building with some fairly different implications for how we even use it.
Today though, we are talking about a new class of models which you might refer to as AI judgment models.
Rather than producing long strings of text, these judgment models, like the one we're discussing today, Jev from TypeSafe, produce probabilities around specific questions.
Is this customer angry?
Is there a new dependency in this email?
Do we need to change the operational plan because of this?
Today, we're exploring the idea behind these models, how they're trained differently, how they can produce these judgments much more quickly and much less expensively, and most importantly, where they're going to fit in your overall model stack.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in.
First of all, thank you to today's sponsors, KPMG, Blitzy, Section, and HyperAgent.
To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts.
Subscriptions are just $3 a month for ad-free.
And if you want to learn about sponsoring the show, send us a note at sponsors at aidailybrief.ai, or just go to aidailybrief.ai, where you can learn all about it.
The AI safety discourse continues to trickle out through the tech industry as well as mainstream society.
But for now, unless something absolutely seismic happens, we're going to move it into the headlines and away from the main episode.
With that in mind, after staying quiet over the weekend, Mark Zuckerberg has made his thoughts known on this idea of an AI slowdown.
On Tuesday, Zuckerberg wrote in a post on X, Every lab has the responsibility and incentive to move at the pace required to train its model safely, and the ability to take its own actions to ensure that happens.
Basically, his view is that pacing is the responsibility of individual labs, rather than a collective action.
And that view hinges on two core ideas.
First, that quote, People don't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned.
And two, labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well.
Emphasizing the point, Zuckerberg said that Meta had delayed the release of Muse by several months to work on safety.
He continued, We didn't call for everyone else to do this before we would.
We just did it as part of our day-to-day work because it was clearly the right thing for people and for us.
Essentially, Zuckerberg is saying that the individual incentives and consequences that are already in place are enough to force AI labs to work on alignment and to pace the frontier correctly, rather than needing some exogenous government-enforced slowdown.
Now, Zuckerberg did support the idea of independent evaluators and advisors as a matter of best practice rather than regulation.
He claimed that Meta has already engaged outside evaluators, not because it's required of them, but because it helps produce better work.
Finally, he concluded, Committing the significant majority of compute towards serving people, rather than racing towards recursive self-improvement, is one of the best ways to ensure that we develop this technology safely.
Meta has made this commitment, and other labs can do this as well.
I believe the key to building a positive future for everyone is maintaining the right balance of power.
This is within our power to do.
It's carefully worded, but basically this whole thing says, come on guys, let's please stop with the theatrics.
And a lot of people frankly found this a breath of fresh air.
YouTuber Joseph Carlson wrote, Hold up a minute, you are telling me companies can slow down, make sure things are safe, without telling all their competitors to slow down?
And I would say that broadly speaking, reactions fell into one of two categories.
The first was like that one and like this from Matthew Berman, love this, model safety and alignment is a feature and economically incentivized.
Or Bill Ackman, who simply called this the proper approach to AI development.
But then on the other side were those, arguing effectively that Zuckerberg just does not have the trust or standing to make this argument.
regardless of the merits of the argument itself.
Hard Fork's Kevin Roos wrote, Whether you're a Democrat or a Republican, an EA or an EACC, I think we can all agree that the person best suited to protect us against the harms of powerful new technology is Mark Zuckerberg.
Now, meanwhile, over in Strange Bedfellows Daily, Bernie Sanders and Steve Bannon have joined forces to call for human-centric AI regulations in the strangest alliance of this political cycle.
The two highly ideological leaders spoke from the same stage on Tuesday at the Future of Life Institute's Pro-Human Assembly in Washington.
Sanders told the crowd, if the lives of every man, woman, and child are going to be fundamentally changed by this technology, then the people of this country must make decisions about AI and not just a handful of oligarchs.
Bannon had very similar remarks, stating, the American citizens are not going to be supplicants to the oligarchs anymore.
We can't do it.
This is a hinge in history.
We have to handle this correctly.
To handle it correctly, number one, we can never trust what an oligarch says.
Now, I will note that while this seems strange at first, these two highly ideologically opposed people sharing the same stage, the broader movements they connect to have a fair bit of shared context and history.
The 2016 presidential election, in which Bernie was narrowly beaten by Hillary to miss out on becoming the Democrat nominee, and in which Bannon obviously architected the first Trump administration, both had their roots in populist anger at the post-GFC financial landscape, and the lack of accountability for the institutions and institutional leaders who were involved in creating that particular economic crisis.
Obviously, the left and the right's reaction to that particular context were different, but it is, I would contend, perhaps less surprising than you think, that at some point these two would find common ground.
And to be clear, I am not dismissing the common ground that they find just on the merits of this particular issue itself, and the power of this particular issue to scramble existing political alliances.
I'm just making the point that their stories are actually more intertwined than you might think at first glance.
In any case, the event seemed somewhat less about the existential risks of AI that have filled the headlines this week, and more about the class struggle the technology has come to represent.
TV screens played a parody interview from a fictional AI CEO, described as someone who loves people but isn't crazy about humans.
Throughout the event, it seemed that AI itself wasn't the risk that needs addressing, but rather the unchecked power of oligarchs.
Now, you might remember that back in August, journalist Jasmine Sun towards the country speaking to real people who were working in opposition to data centers, and found that the most common complaint was not about electricity, water use, or noise pollution, but a lack of control in the sense that the future was being forced upon their communities with little input.
The same was true about this event.
Across multiple speakers, the risk of AI was not framed around cybersecurity, bioterrorism, or other X-risks.
It was about a lack of agency in determining the shape of the future.
From this followed their main concern, which was simply the idea of losing control of AI.
As Texas Democrat Greg Kassar, who is sponsoring Bernie Sanders' superintelligence bill, said, the answer is simple.
We ban AI systems that are too powerful for humans to control.
And as strange as it might seem, despite all the intensity of this rhetoric recently, I still contend that we've moved into a new phase that's all about negotiating the relationship where citizens and governments have a stake in this, and where given that that is now pretty much where everyone is, the next phase is likely to include a lot more specificity and, dare I say, nuance.
Glenn Beck, for example, had a long monologue on his show about how he can both have signed the pro-human AI declaration alongside them, but also disagree with them on specifics of data centers.
And honestly, as crazy as it sounds, I think that the more that the political discourse kind of frags your brain for how confusing and all over the place it is, that might just be a sign that we're actually making progress.
The conversation also found its way into Salesforce's annual Dreamforce event, where the company unveiled a new AI model.
but where a lot of the chatter on social media focused on appearances from Sam Altman, Jensen Huang, and Dario Amadei, who each iterated their own safety views.
Dario argued that he was just trying to put forward a set of standards that the industry could organize around, while Altman said that he was very confident in our company's ability and our industry's ability to do this safely.
Jensen Huang, meanwhile, reiterated the Zuckerbergian view, saying, run as fast as you can, but if you feel at any given point in time the company's out of control or the product's not going to be safe, take a pause and make sure you get it right.
And as for Salesforce CEO Mark Benioff, he believes that every company has a responsibility to uphold ethical standards.
Still, to give the safety debate a bit of a rest, Salesforce also had two big practical AI announcements.
First, they're releasing their first in-house model in quite some time.
Called CoA, the model is a fine-tune of NVIDIA's Nemotron and is designed to handle sales management within the CRM.
The announcement reinforces the role that open source has to play in the enterprise by enabling this sort of narrow vertical model.
Secondly, Salesforce unveiled a new initiative called AI Force.
This will be the umbrella term for Salesforce's connectors that will allow third-party agents to access Salesforce data.
The release reinforces Salesforce's commitment to moving towards headless software in a platform-agnostic way, allowing any agent to become the interface.
Now, obviously, these are big conversations that are important and are going to continue, but for now, that is where we will close the headlines.
Next up, the main episode.
A new study from KPMG in the University of Texas at Austin found that when people work with AI, Similar skills don't guarantee similar outcomes.
Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs.
These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI.
Learn more about what separates AI amplifiers from everyone else at kpmg.com slash us slash AI amplifiers.
Here's why most legacy modernization projects fail.
The AI doing the work can't understand code bases at scale.
It sees a small slice of context, examines syntax, and misses years of decisions distributed across the global application ecosystem.
Blitzy solves this the way it solves everything.
Grounded in your code before any migration begins, Blitzy's agents reverse engineer the entire legacy system into a persistent knowledge graph.
Every dependency, every constraint, every piece of tribal knowledge that used to live in one engineer's head.
From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith-to-microservices transformations, all validated end-to-end.
One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137-week baseline with coding agents.
That's 9x compression.
Retire technical debt while accelerating your roadmap.
See how at Blitzy.com.
That's B-L-I-T-Z-Y dot com.
Here's a harsh truth.
Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized.
Half of companies have AI tools, but only 12% use them for business value.
Most employees are still using AI to summarize meeting notes.
If you're the one responsible for AI adoption at your company, you need Section.
Section is a platform that helps you manage AI transformation across your entire organization.
It coaches employees on real use cases, tracks who's using AI for business impact, and shows you exactly where AI is and isn't creating value.
The result?
You go from rolling out tools to driving measurable AI value.
Your employees move from meeting summaries to solving actual business problems.
And you can prove the ROI.
Stop guessing if your AI investment is working.
Check out Section at sectionai.com.
That's S-E-C-T-I-O-N-A-I dot com.
This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.
Forget local agents and chat workflows waiting on your laptop to be prompted.
Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses.
Marketing agents turn competitor moves into landing pages.
Sales agents enrich leads, draft emails, and updates the CRM.
Ops agent chases the paperwork and tracks the budget.
Every agent has access to shared context and follows your rules about scope and approvals.
It's time you add agents that feel like teammates.
Hire yours at Hyperagent.
Get $100 in credits at hyperagent.com slash AI Daily Brief.
Welcome back to the AI Daily Brief.
On today's main episode, we are looking at something really different and quite rare, which is, in short, a totally different approach to AI that is not just another LLM.
Now, in the nearly four years since ChatGPT was released and the three and a half that the show has been around, the vast majority of the things that we and the rest of the industry have focused on have been in some ways related to large language models.
And yet, as some have pointed out, Although often in a way that didn't get much traction and was more or less screaming into the void, LLMs are not, in fact, the totality of artificial intelligence.
Yesterday, Diogo Almeida posted on X, After co-inventing ChatGPT, I kept asking myself, why have superhuman chat models not led to AGI?
I've spent the last two years in stealth building a new way to train models, RLCD, or reinforcement learning for calibrated decisions, and a new type of frontier AI model that we are releasing today.
JEV.
JEV is 20 to 200 times faster, 40 to 400 times cheaper, with output tokens free, frontier composable intelligence optimized for decisions.
As far as I can tell, the shortest path to AI-based economic revolution.
Now, one could absolutely be forgiven for seeing numbers like 20 to 200 times faster and 40 to 400 times cheaper, and being a bit skeptical of the claims to say the least.
And were this just another LLM, that skepticism would be entirely warranted.
However, with JEV and this new strategy, we're dealing with something that is quite different.
The company behind JEV is called TypeSafe.
And in the announcement blog post, they talk a little bit more about what makes their approach different.
They write, We built a new stack entirely focused on automation, with a new model architecture, parallel sampler for maximum efficiency, and training method we call reinforcement learning for calibrated decisions.
Whereas existing LLMs optimize for human preference, i.e.
write-ups and chat responses that human raters prefer, the new System 1 models, the first of which is JEV, optimize for calibrated decisions, or answers with epistemically honest probabilities.
So what does that actually mean?
Well, let's look at how Mike Taylor from Every describes it.
He writes, Think of it as a smart if-then statement that determines what happens next when you're automating a workflow.
Say you're building software that prioritizes customer service requests, and you write code that asks the model, does this customer sound angry?
Jev might answer 0.9, which means there's an estimated 90% probability that the answer is yes based on what the model learned in training.
You could also provide categories you define like annoyed, irritated, offended, furious, and enraged, and learn that the customer was 60% likely to be classified as furious, with only a 10% probability of being enraged.
Going on to explain why this matters and how it differs than LLMs, Mike continues, with an answer of 0.9, very likely to be angry, the software might automatically proceed to escalate the customer concern to a manager.
Or if it answers 0.1, not likely to be angry, that request might be deprioritized.
However, chatbots are trained to respond with flowery text, like, you're absolutely right, this customer does sound very angry.
Would you like me to compose a draft email response in a friendly, supportive tone?
This text response would cause the program you're building to crash because it was expecting a number between 0 and 1, not an essay.
Teal fellow Michael Lee says this allows a class of decision-making that was neither suited to dumb, unintelligent code nor to slow, expensive LLMs.
As Chubby sums up, JEV is an AI model built for decisions rather than text generation.
And it is important to note that the trade-off here is that this model does not generate text.
It is not, in other words, a replacement.
for LLMs in general.
It's a replacement for a certain category of work that LLMs do, where they've been very square peg mashed into a round hole to do it.
The idea, continues Chubby, is to embed fast, cheap AI decisions into software.
So what is this model actually built for?
Well, think of how much of office work consists of reading something and deciding what should happen next.
Does this message need a response?
Which department should handle it?
Does this document answer the question?
Is this customer describing a bug or asking for a feature?
Does this draft make a claim its source doesn't support?
Is the situation routine enough to automate or should someone review it?
These are judgments about meaning, and they're often difficult to express as fixed rules.
JEV then is designed to take the relevant information and answer narrowly defined questions with probabilities, categories, or scores.
The surrounding software then uses those answers to route, rank, flag, or proceed.
In the TypeSafe documentation, they explicitly recommend breaking complex decisions into small questions and then combining their results in code.
So where would this show up in normal business?
One obvious area is customer support, where the small judgment the model could make would be something like, is the customer frustrated and have previous replies failed to address it?
With those small judgments in hand, the software could then next route the ticket, raise its priority, or request human review.
In the sales domain, the model could judge, is this a buying inquiry?
Does the prospect fit the product?
Are they requesting a meeting?
The software could then take those judgments to sort inbound leads and assign follow-up.
In marketing and editorial, the model might judge whether the copy meets specific style rules or whether the offer is clear.
The software that surrounds it could then flag passages for revision before publication.
And importantly, where a lot of people went was not just understanding where they would use this instead of LLMs, but how they might use it alongside LLMs.
YC founder Nathan Flurry wrote, I'd imagine a lot of workflows that look like LLM proposes options, JEV decides, code executes.
To put a clear example on this, imagine that a customer writes, this is the third time I've contacted you, we still can't export our reports, and our renewal is next week.
A support workflow supported by something like JEV, could ask several questions together.
One, is the customer describing a product problem?
Two, does the message indicate repeated unsuccessful support?
Three, is a commercially significant deadline approaching?
Four, which team is best equipped to help?
The software that surrounds it could then combine those signals with actual account information, such as the renewal date, and escalate the ticket.
From there, you would still have a generative model draft the reply, but Jev once again could check the draft against narrow criteria.
Does the response acknowledge the repeated contacts?
Does it address the export problem?
Does it promise something unsupported by the information provided?
And part of what people are excited about opening up with this new approach is that cheap judgment makes frequent checking more practical.
If a check adds a noticeable delay or expense, a team may run it only on selected cases or at the end of a task.
If it becomes sufficiently fast and inexpensive, It could run on every incoming request after each draft revision across many candidate documents before an agent takes a consequential step.
Back in Mike Taylor's article from Every, he gave Jev the text from all 27 of his articles alongside 10 deliberately AI-styled counterpoints, then asked the same 21 questions concurrently across all articles to check for AI tells.
Basically, the check that he was doing with Jev was does this essay do specific things that indicate to people that it is AI composed?
Things like, does the text repeat an idea without adding evidence?
Does it force a symmetrical both sides argument?
Does it over-explain a straightforward point?
In less than 0.7 seconds, Mike said, Jev quote-unquote read all 37 documents and answered all 21 questions for each, returning 777 judgments for an estimated quarter of a cent.
As he points out, that's fast and cheap enough to AI check everything everyone at your company has ever written and get the results back in an instant.
A comparison that Mike makes is a code linter for knowledge work.
He writes, Perengrat says, Most software is ultimately a giant tree of if this, do that, if this, root here, if this, escalate, if this, reject, if this, ask a human.
Jev is basically asking, What if those if statements could understand messy human context?
That's a much more interesting framing than another AI model.
I can see this being very useful for fraud and risk, support routing, moderation, PR and QA automation, lead scoring, compliance, workflow orchestration, and agent routing.
Early tech, obviously, he says, but the category itself makes a lot of sense.
Now, interestingly, Matt Stockton points out that in some ways, companies adopting this amounts to a post-LLM AI technology, making pre-LLM machine learning techniques a little bit more accessible.
As he writes, With the emergence in popularity of LLMs, more companies are thinking, maybe we can use AI for that, and are solving classification and regression problems with LLMs.
This is good in some ways because companies are potentially automating some manual work, but also bad in some ways because it's often the wrong tool for the job, and possibly not as good as classic ML methods for what they are trying to do.
But he points out, the classical techniques require you to label your data, train a model, and host that model somewhere.
They aren't as easy to use compared to calling an LLM API.
And it requires you and your org to be aware of those techniques and capable of investing in them.
Without going too far on the analogy, he basically says one way to look at Jev is as a UX for using LLM-style user interaction patterns for classical ML techniques.
Now, one interesting question that comes up is whether this is for individuals or teams building systems.
And the short answer is that while it is absolutely both, it also puts a fine point on the multiplayer AI themes that we've been talking about recently.
Certainly, individuals could use this sort of capability to have personal tools that sort an inbox against their own priorities, check drafts of their writing against an editorial rubric, rank saved articles against research interests, or flag commitments in meeting transcripts, things like that, that are going to personally help you do your work better.
But I think that where this sort of technique is going to really shine is in the domain of teamwork that happens through small judgments about who needs to know, who should act, and whose approval is required.
When an agent is serving a single person, it gets pretty far simply by learning that person's preferences.
An agent operating across a team, however, needs to understand the relationships between people's work.
That creates a different set of questions.
Who owns this?
Whose work does this affect?
Is someone waiting on this decision?
Does this promise create an obligation for another team?
Can the current owner decide or does this need broader agreement?
In some ways, the interesting unit of work becomes the handoff.
Consider a salesperson telling a customer, we should be able to support that integration before your renewal.
For the salesperson's personal agent, the next steps might be straightforward.
Update the account record, draft a follow-up, and create a reminder.
But inside the organization, that sentence implicates several responsibilities.
For the sales folks, that sentence means for them a potential way to secure the renewal, i.e.
supporting that integration.
For engineering, that sentence means a possible delivery commitment involving uncertain work.
For product, it means a potential change to roadmap priorities.
And for customer success, it's an expectation they may have to manage.
A multiplayer agent would need to recognize the sentence as a possible cross-team commitment, and a judgment model could assess specific questions.
Does this message imply a delivery promise?
Does the promise concern work outside the speaker's authority?
Does it conflict with the supplied roadmap?
Is there evidence that the responsible team agreed?
The larger system could then create a proposed commitment, identify the necessary owners, and request the missing decisions.
Now again, to reinforce costs and benefits.
The big cost of JEV and this type of judgment model in general to the extent that this becomes a category, is that it is an incomplete category by definition.
It cannot do all the work that we currently have generative AIs do.
Judgment models are going to have to be part of a more complex model architecture, the type of model stack that we've been discussing for the last several months.
The benefit, of course, though, is that it can do this extraordinarily inexpensively.
And because this sort of judgment intelligence can be applied so cheaply, it means that done well, it can be integrated incredibly deeply into the automated systems we're all building.
This is obviously just the first day of a very new concept, and a concept which is significant enough that this very competent team has spent two years working on, so obviously we're going to need to see how it all plays out in practice.
But it does have the feel when you dig in of something both important and obvious, the type of thing that once it exists, we will be surprised in the future that we didn't have it for so long.
Certainly, I'm going to be keeping an eye on this and I will continue to look out for more examples of how people are using it, as well as where companies are running into challenges as they try to build these new types of systems.
For now, though, very cool stuff to go check out.
And that's going to do it for today's AI Daily Brief.
Appreciate you listening or watching as always.
Until next time, peace.
