← Reinforcement Learning

What podcasts say about Reinforcement Learning

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about Reinforcement Learning

7 statements · 1 positive · 2 negative · 2 mixed · 2 neutral

  1. Noam BrownNegativeSep 17, 2026· Dwarkesh Podcast

    AI progress could slow if challenging reinforcement-learning problems become scarce.

    “if we run out of problems to ask it that challenge it, then that is a plausible scenario where actually like, okay, it becomes much harder to make progress”

    Listen at 8:12

    Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvement
  2. Beren MillidgeMixedSep 11, 2026· Dwarkesh Podcast

    Much apparent RL progress actually comes from strong synthetic mid-training data.

    “An awful lot of what we see as successes of RL actually comes from very, very good mid-training data”

    Listen at 1:18:57

    Open the episode · AI researchers debate how close we are to recursive self-improvement
  3. LLM reinforcement learning has produced horizon generalization more than broad cross-domain reasoning transfer.

    “what we did get though is horizon generalization”

    Listen at 1:22:52

    Open the episode · AI researchers debate how close we are to recursive self-improvement
  4. Ajeya CotraNeutralSep 1, 2026· Dwarkesh Podcast

    Reinforcement learning creates software capable of creatively pursuing goals.

    “RL, the whole point of RL is to create goal-oriented beings, software that can creatively pursue goals.”

    Listen at 53:59

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
  5. Dwarkesh PatelPositiveJun 26, 2026· Dwarkesh Podcast

    Reinforcement learning concentrates model updates on information relevant to achieving outcomes.

    “RL is great at concentrating the update to only what is relevant to getting the outcome right.”

    Listen at 14:32

    Open the episode · The next big breakthrough will be AIs learning on the job
  6. Dwarkesh PatelNeutralJun 19, 2026· Dwarkesh Podcast

    Reinforcement learning requires models to have prior probability for correct solutions.

    “For this process to work, the model must have at least some prior probability to anticipate the correct solution in the first place.”

    Listen at 0:47

    Open the episode · The data black hole at the center of AI
  7. Andrej KarpathyNegativeOct 29, 2022· Lex Fridman Podcast

    Reinforcement learning is extremely inefficient for training neural networks.

    “reinforcement learning is an extremely inefficient way of training neural networks”

    Listen at 54:26

    Open the episode · #333 – Andrej Karpathy: Tesla AI, Self-Driving, Optimus, Aliens, and AGI

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Reinforcement Learning: what podcasts say · PodLume