← Reinforcement learning from verifiable outcomes

What podcasts say about Reinforcement learning from verifiable outcomes

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about Reinforcement learning from verifiable outcomes

3 statements · 1 positive · 2 negative

  1. RLVR scales by repeatedly testing models on increasingly difficult verifiable problems.

    “with RLVR you literally give the model well, you let the model solve more and more complex, difficult problems.”

    Listen at 2:08:45

    Open the episode · #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
  2. Dwarkesh PatelNegativeDec 23, 2025· Dwarkesh Podcast

    Training on verifiable outcomes is doomed if human-like learners are near.

    “If we're actually close to a human-like learner, then this whole approach of training on verifiable outcomes is doomed.”

    Listen at 0:07

    Open the episode · An audio version of my blog post, Thoughts on AI progress (Dec 2025)
  3. Dwarkesh PatelNegativeDec 23, 2025· Dwarkesh Podcast

    No well-fitting public scaling trend is known for reinforcement learning from verifiable reward.

    “for which we have no well-fit publicly known trend.”

    Listen at 8:51

    Open the episode · An audio version of my blog post, Thoughts on AI progress (Dec 2025)

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Reinforcement learning from verifiable outcomes: what podcasts say · PodLume