← Off-policy reinforcement learning

What podcasts say about Off-policy reinforcement learning

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about Off-policy reinforcement learning

1 statement · 1 negative

  1. Eric JangNegativeMay 15, 2026· Dwarkesh Podcast

    Training on unreachable off-policy states wastes model capacity.

    “if the current model is looking at states that it would never reach, then it's kind of wasting capacity”

    Listen at 2:10:51

    Open the episode · Eric Jang – Building AlphaGo from scratch

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Off-policy reinforcement learning: what podcasts say · PodLume