What podcasts say about AI model evaluations
Every statement, with the speaker, the exact quote and the moment it was said.
What experts have said about AI model evaluations
6 statements · 2 positive · 3 negative · 1 neutral
Accenture will evaluate models, test safeguards, and assess alignment with human preferences.
“So Accenture will be used to evaluate the newest models, test safeguards, and ensure the technology is aligned with what humans actually want.”
Open the episode · Novo Declines; Critical Metals Soars; WBD HigherListen at 3:43
Short release cycles may prevent full evaluation of long-horizon AI capabilities.
“If you're in a world where they can operate effectively over 3 months, but the model release cycle is every 2 months, then you don't have a way to evaluate the models at the full length of their capabilities before the model release cycle”
Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvementListen at 1:03:57
Realistic evaluation environments could help predict real-world AI behavior.
“if you can create very realistic environments and put the AIs in there. If you have a sufficiently realistic evaluation environment, then you can get a sense of, okay, is the AI actually going to behave well when we deploy it in the real world?”
Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvementListen at 1:16:02
Creating evaluations indistinguishable from reality is becoming increasingly difficult.
“making an environment that's realistic enough that it matches, that it's indistinguishable from the real world for them is becoming increasingly more difficult”
Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvementListen at 1:17:13
AI companies may evaluate one another’s models or use independent evaluators
“potentially even grading each other's models or having independent evaluators do that”
Open the episode · Meta CEO Weighs In On AI Safety DebateListen at 30:18
Without evaluations, teams may not detect emergent capability jumps.
“unless you have the evals, unless you have the systems to test these jumps might actually happen. And you don't know.”
Open the episode · Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne PennListen at 19:07
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.