
Sep 16, 2026 · 25 min
Listen from 0:06
Listen at 0:06
Judgment models push AI beyond text generation
Why a New Class of AI “Judgment Models” Could Have Big Business Implications
Fast probabilistic assessments could give AI agents a cheaper way to verify work and coordinate consequential business decisions.
- 1Jev represents an AI model class that answers targeted questions with probabilities rather than generating long-form text.
- 2Judgment models could help agents check their own work and support coordination across business processes.
- 3The episode places Jev alongside debates over AI regulation, Zuckerberg’s opposition to a collective slowdown, and Salesforce’s expanding agent ecosystem.
Don't miss
Whittemore’s discussion of Jev frames probabilistic judgment—not text generation—as a possible foundation for agents that verify their own work.
The brief
Nathaniel Whittemore introduces judgment models, a category aimed at answering specific questions with probabilities instead of producing long-form prose.
The episode focuses on Jev, a model from Typesafe designed to deliver fast, inexpensive assessments that could make AI systems more useful in practical settings.
The central business question is whether agents can use judgment models to verify their work, assess uncertainty, and coordinate decisions without relying on another lengthy generation.
Jev’s implications sit within a broader AI landscape shaped by regulation debates, Zuckerberg’s opposition to a collective slowdown, and Salesforce’s expanding agent ecosystem.
The standout idea is that AI progress may increasingly depend on compact systems that evaluate options and confidence, not just models that produce more text.
What was said on this episode
18 statements · 15 positive · 2 negative · 1 mixed
Typesafe’s Jev produces probabilities for specific questions rather than long-form text.
“these judgment models like the one we're discussing today, Jev from Typesafe, produce probabilities around specific questions”
Listen at 0:18
Judgment models can produce judgments faster and more cheaply than conventional approaches.
“how they can produce these judgments much more quickly and much less expensively”
Listen at 0:35
Model-related harm creates significant liability for AI labs.
“labs face significant liability if their models cause harm”
Listen at 2:07
Prioritizing compute for serving users over recursive self-improvement improves AI safety.
“committing the significant majority of compute towards serving people, rather than racing towards recursive self-improvement, is one of the best ways to ensure that we develop this technology safely”
Listen at 2:56
AI policy debate will likely become more specific and nuanced.
“the next phase is likely to include a lot more specificity, and dare I say nuance”
Listen at 7:04
Increasingly confusing AI political discourse may indicate progress.
“the more that the political discourse kind of frags your brain for how confusing and all over the place it is, that might just be a sign that we're actually making progress”
Listen at 7:21
Salesforce’s Koa is designed for sales management within CRM systems.
“the model is a fine-tune of NVIDIA's Nimotron and is designed to handle sales management within the CRM”
Listen at 8:19
Salesforce AI Force enables third-party agents to serve as software interfaces.
“allowing any agent to become the interface”
Listen at 8:44
Top AI users increase value by guiding, evaluating, and refining outputs.
“the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs”
Listen at 9:10
Typesafe claims Jev is 20–200 times faster and 40–400 times cheaper than alternatives.
“Jev is 20 to 200 times faster, 40 to 400 times cheaper with output tokens free, frontier composable intelligence optimized for decisions”
Listen at 12:52
Jev processed 37 documents and 777 judgments in under 0.7 seconds for about $0.0025.
“In less than 0.7 seconds, Mike said, Jev quote unquote read all 37 documents and answered all 21 questions for each. Returning 777 judgments for an estimated quarter of a cent.”
Listen at 19:18
Jev could function like a code linter for knowledge work.
“it could act as a kind of code linter for knowledge work”
Listen at 19:57
Jev can apply conditional workflow logic to messy human context.
“Jev is basically asking, what if those if statements could understand messy human context?”
Listen at 20:16
Judgment models could be useful across fraud, support, moderation, compliance, and workflow tasks.
“I can see this being very useful for fraud and risk, support routing, moderation, PR and QA automation, lead scoring, compliance, workflow orchestration, and agent routing”
Listen at 20:24
Judgment models cannot perform all tasks currently handled by generative AI.
“The big cost of JEV and this type of judgment model in general, to the extent that this becomes a category, is that it is an incomplete category by definition.”
Listen at 23:48
Judgment models can perform their specialized tasks at very low cost.
“The benefit, of course, though, is that it can do this extraordinarily inexpensively.”
Listen at 24:10
Low-cost judgment intelligence can be deeply integrated into automated systems.
“it means that done well, it can be integrated incredibly deeply into the automated systems we're all building”
Listen at 24:18
Judgment models appear important and conceptually obvious after examination.
“it does have the feel when you dig in of something both important and obvious”
Listen at 24:34
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.