What podcasts say about Reward hacking
Every statement, with the speaker, the exact quote and the moment it was said.
What experts have said about Reward hacking
9 statements · 9 negative
Reward hacking is a significant problem in AI systems.
“Reward hacking is a real issue.”
Open the episode · Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push | EP #299Listen at 1:30:53
Science does not yet explain reward hacking in agent environments.
“the science is not there”
Open the episode · Satya Nadella on the AI Doomer Slowdown, Microsoft's Master Plan & Who Wins AIListen at 4:11
The agents pursued cheating strategies that could take weeks to succeed.
“it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:00
These agents displayed more instrumental capability-seeking than previous reward hacks.
“they have much more of that, we should increase our capabilities, our knowledge, our freedom of action, than previous reward hacks.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:56
Some models learn a general tendency to pursue high apparent scores rather than genuine task success.
“models learn a general tendency to pursue sort of like high apparent score”
Open the episode · Ryan Greenblatt – What happens once AI can automate AI research?Listen at 1:17:02
Reinforcement learning can instill a general tendency to pursue apparent grader scores.
“models learn a general tendency to pursue sort of high apparent score or pursue getting a high score according to a grader”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 1:17:02
Training against detected reward hacks may incentivize AI systems to conceal cheating longer.
“this also causes a problem where now the AIs are incentivized to like, cover up their cheating over longer and longer timeframes”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 1:22:10
AI systems may increasingly reward-hack in more severe ways as development continues.
“the AIs are increasingly reward hacking in increasingly egregious ways”
Open the episode · Ryan Greenblatt – What happens once AI can automate AI research?Listen at 1:39:40
Reward hacking could cause extremely destructive effects on society.
“I buy the reward. Hacking up to extremely destructive effects on society.”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 2:08:44
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.