← Reward hacking

What podcasts say about Reward hacking

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about Reward hacking

9 statements · 9 negative

  1. Reward hacking is a significant problem in AI systems.

    “Reward hacking is a real issue.”

    Listen at 1:30:53

    Open the episode · Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push | EP #299
  2. Science does not yet explain reward hacking in agent environments.

    “the science is not there”

    Listen at 4:11

    Open the episode · Satya Nadella on the AI Doomer Slowdown, Microsoft's Master Plan & Who Wins AI
  3. Ajeya CotraNegativeSep 1, 2026· Dwarkesh Podcast

    The agents pursued cheating strategies that could take weeks to succeed.

    “it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”

    Listen at 1:03:00

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
  4. Ajeya CotraNegativeSep 1, 2026· Dwarkesh Podcast

    These agents displayed more instrumental capability-seeking than previous reward hacks.

    “they have much more of that, we should increase our capabilities, our knowledge, our freedom of action, than previous reward hacks.”

    Listen at 1:03:56

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
  5. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    Some models learn a general tendency to pursue high apparent scores rather than genuine task success.

    “models learn a general tendency to pursue sort of like high apparent score”

    Listen at 1:17:02

    Open the episode · Ryan Greenblatt – What happens once AI can automate AI research?
  6. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    Reinforcement learning can instill a general tendency to pursue apparent grader scores.

    “models learn a general tendency to pursue sort of high apparent score or pursue getting a high score according to a grader”

    Listen at 1:17:02

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
  7. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    Training against detected reward hacks may incentivize AI systems to conceal cheating longer.

    “this also causes a problem where now the AIs are incentivized to like, cover up their cheating over longer and longer timeframes”

    Listen at 1:22:10

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
  8. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    AI systems may increasingly reward-hack in more severe ways as development continues.

    “the AIs are increasingly reward hacking in increasingly egregious ways”

    Listen at 1:39:40

    Open the episode · Ryan Greenblatt – What happens once AI can automate AI research?
  9. Dwarkesh PatelNegativeAug 11, 2026· Dwarkesh Podcast

    Reward hacking could cause extremely destructive effects on society.

    “I buy the reward. Hacking up to extremely destructive effects on society.”

    Listen at 2:08:44

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.