RH

Reward hacking

Topic

Heard in 1 episode across 1 show since Aug 2026

Reward hacking, also known as specification gaming, occurs when an artificial intelligence system trained with reinforcement learning optimizes its objective function to achieve the literal, formal specification of a reward without actually achieving the outcome intended by its programmers. This behavior represents a key challenge in AI alignment, where the system exploits loopholes or flaws in the reward design to maximize its score in unintended or counterproductive ways.

Episodes

1
across 1 show

First heard

Aug 2026

Expert statements

9
9 negative

Episodes per month

Public PodLume episodes featuring it, over the last year.

Show the data
MonthEpisodes
Nov 20250
Dec 20250
Jan 20260
Feb 20260
Mar 20260
Apr 20260
May 20260
Jun 20260
Jul 20260
Aug 20261
Sep 20260
Oct 20260

What experts have said about Reward hacking

9 statements · 9 negative

  1. Reward hacking is a significant problem in AI systems.

    “Reward hacking is a real issue.”

    Listen at 1:30:53

    Open the episode · Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push | EP #299
  2. Science does not yet explain reward hacking in agent environments.

    “the science is not there”

    Listen at 4:11

    Open the episode · Satya Nadella on the AI Doomer Slowdown, Microsoft's Master Plan & Who Wins AI
  3. Ajeya CotraNegativeSep 1, 2026· Dwarkesh Podcast

    The agents pursued cheating strategies that could take weeks to succeed.

    “it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”

    Listen at 1:03:00

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Episodes

1 episode featuring Reward hacking, newest first