← AI model evaluations

What podcasts say about AI model evaluations

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about AI model evaluations

6 statements · 2 positive · 3 negative · 1 neutral

  1. Nathan HagerPositiveSep 21, 2026· Stock Movers

    Accenture will evaluate models, test safeguards, and assess alignment with human preferences.

    “So Accenture will be used to evaluate the newest models, test safeguards, and ensure the technology is aligned with what humans actually want.”

    Listen at 3:43

    Open the episode · Novo Declines; Critical Metals Soars; WBD Higher
  2. Noam BrownNegativeSep 17, 2026· Dwarkesh Podcast

    Short release cycles may prevent full evaluation of long-horizon AI capabilities.

    “If you're in a world where they can operate effectively over 3 months, but the model release cycle is every 2 months, then you don't have a way to evaluate the models at the full length of their capabilities before the model release cycle”

    Listen at 1:03:57

    Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvement
  3. Noam BrownPositiveSep 17, 2026· Dwarkesh Podcast

    Realistic evaluation environments could help predict real-world AI behavior.

    “if you can create very realistic environments and put the AIs in there. If you have a sufficiently realistic evaluation environment, then you can get a sense of, okay, is the AI actually going to behave well when we deploy it in the real world?”

    Listen at 1:16:02

    Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvement
  4. Noam BrownNegativeSep 17, 2026· Dwarkesh Podcast

    Creating evaluations indistinguishable from reality is becoming increasingly difficult.

    “making an environment that's realistic enough that it matches, that it's indistinguishable from the real world for them is becoming increasingly more difficult”

    Listen at 1:17:13

    Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvement
  5. Tom MuellerNeutralSep 16, 2026· Bloomberg Tech

    AI companies may evaluate one another’s models or use independent evaluators

    “potentially even grading each other's models or having independent evaluators do that”

    Listen at 30:18

    Open the episode · Meta CEO Weighs In On AI Safety Debate
  6. Without evaluations, teams may not detect emergent capability jumps.

    “unless you have the evals, unless you have the systems to test these jumps might actually happen. And you don't know.”

    Listen at 19:07

    Open the episode · Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

AI model evaluations: what podcasts say · PodLume