What podcasts say about Reinforcement Learning
Every statement, with the speaker, the exact quote and the moment it was said.
What experts have said about Reinforcement Learning
7 statements · 1 positive · 2 negative · 2 mixed · 2 neutral
AI progress could slow if challenging reinforcement-learning problems become scarce.
“if we run out of problems to ask it that challenge it, then that is a plausible scenario where actually like, okay, it becomes much harder to make progress”
Open the episode · Noam Brown – Agent swarms, alignment, & recursive self-improvementListen at 8:12
Much apparent RL progress actually comes from strong synthetic mid-training data.
“An awful lot of what we see as successes of RL actually comes from very, very good mid-training data”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:18:57
LLM reinforcement learning has produced horizon generalization more than broad cross-domain reasoning transfer.
“what we did get though is horizon generalization”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:22:52
Reinforcement learning creates software capable of creatively pursuing goals.
“RL, the whole point of RL is to create goal-oriented beings, software that can creatively pursue goals.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 53:59
Reinforcement learning concentrates model updates on information relevant to achieving outcomes.
“RL is great at concentrating the update to only what is relevant to getting the outcome right.”
Open the episode · The next big breakthrough will be AIs learning on the jobListen at 14:32
Reinforcement learning requires models to have prior probability for correct solutions.
“For this process to work, the model must have at least some prior probability to anticipate the correct solution in the first place.”
Open the episode · The data black hole at the center of AIListen at 0:47
Reinforcement learning is extremely inefficient for training neural networks.
“reinforcement learning is an extremely inefficient way of training neural networks”
Open the episode · #333 – Andrej Karpathy: Tesla AI, Self-Driving, Optimus, Aliens, and AGIListen at 54:26
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.