Dwarkesh Podcast
Dwarkesh Podcast

May 15, 2026 · 2h 38m

AlphaGo architecture reveals how reinforcement learning drives artificial general intelligence

Eric Jang – Building AlphaGo from scratch

Understanding the architectural leap from board game AI to self-improving LLMs reveals how modern systems are beginning to automate their own scientific discovery.

3 key takeaways
  1. 1Monte Carlo Tree Search and self-play provide the foundational primitives necessary for scaling machine intelligence beyond human limits.
  2. 2Inference-time scaling and reinforcement learning are shifting the paradigm of how modern large language models are trained and deployed.
  3. 3Automated LLM loops successfully manage complex hyperparameter optimization but still lack high-level human scientific intuition.

Don't miss

Eric Jang explains how automated LLM loops conduct AI research and where they hit a wall in scientific intuition.

The brief

AI researcher Eric Jang joins Dwarkesh Patel to break down the mechanics of AlphaGo, demonstrating how Monte Carlo Tree Search and self-play serve as fundamental primitives for building intelligence.

The transition from board game AI to modern large language models shows how reinforcement learning and inference-time scaling are actively reshaping the path toward artificial general intelligence.

Jang details his hands-on experience using automated LLM loops to conduct AI research, showcasing how automated systems can already handle complex tasks like hyperparameter optimization.

While automated loops excel at optimization, the frontier of AI development still faces major bottlenecks in replicating high-level human scientific intuition and breakthrough reasoning.

What was said on this episode

35 statements · 26 positive · 6 negative · 3 neutral

  1. Eric Jangon KataGoPositive1:49

    KataGo reduced training compute for strong Go bots by approximately forty times.

    “achieved a 40x reduction in compute needed to train a really strong GoBot”

    Listen at 1:49

  2. LLM coding has reduced AlphaGo-like implementation costs from millions to thousands of dollars.

    “what took a whole team of research scientists at DeepMind and millions of dollars of research and compute can now be done for a few thousand dollars of rented computer”

    Listen at 2:05

  3. Human Go players use an implicit value function to evaluate whether a board position is winnable.

    “humans as implicitly having a neural network called a value function that basically takes in a board state and then it kind of evaluates key win”

    Listen at 25:21

  4. A trained value function can resolve Go positions without exhaustively searching deeply.

    “you can train a value function to look at a board and quickly resolve the game without playing out all of these trees into a very deep search depth”

    Listen at 26:30

  5. Eric Jangon AlphaGoPositive27:47

    AlphaGo makes both Go tree breadth and search depth computationally tractable.

    “AlphaGo gives us a way to basically shrink both of those to be very tractable”

    Listen at 27:47

  6. Residual networks outperform transformers for low-budget Go experiments.

    “my experience is that resnets still kind of outperform transformers and kind of give you more bang for the buck at lower budgets”

    Listen at 33:14

  7. Transformers outperform residual convolutional networks when tasks require more global context.

    “transformers start to outperform residual convolutional networks when you want more global context”

    Listen at 33:28

  8. Transformers require more data to learn invariant local features in vision tasks.

    “you do need more data there so that you can kind of learn through data the sort of invariant, local, local features”

    Listen at 34:40

  9. Perfect-information games have Nash-equilibrium strategies no worse than other strategies.

    “in perfect information games there does exist a Nash equilibrium strategy for which you can do no worse than any other strategy”

    Listen at 36:06

  10. The Nash-equilibrium strategy used by Go agents appears unbeatable by human strategies.

    “The Nash equilibrium seems to be superhuman. No human strategy seems to be able to beat it.”

    Listen at 36:43

  11. A policy network trained on expert games can play Go quickly and strongly without search.

    “if you just take this policy recommendation and take the Argmax over, if this is the probabilities, if you take the Argmax and you just take this action as your go play, it'll be a very, very fast go player that doesn't think in terms of reasoning steps. It just kind of shoots from the hip and it'll be a very strong go player, which is already quite miraculous”

    Listen at 42:11

  12. Eric Jangon Modern Go botsPositive44:29

    Modern Go bots require relatively little test-time compute.

    “modern go bots don't need that much compute at test time”

    Listen at 44:29

  13. Explicit policy modeling improves Monte Carlo tree search feedback and recursive self-improvement.

    “having this as an explicit entity you're modeling rather than an implicit normalization over your value, is a good idea”

    Listen at 57:31

  14. AlphaGo training distills the outcome of search into the neural network policy.

    “just train this to approximate the outcome of 1000 steps of search”

    Listen at 1:05:35

  15. MCTS convergence is guaranteed only in the limit of infinitely many simulations.

    “It's only guaranteed to converge when you kind of take N to infinity.”

    Listen at 1:09:36

  16. Monte Carlo tree search does not always improve the policy network.

    “it's not a guarantee to improve”

    Listen at 1:09:51

  17. Practitioners should first establish a strong value function before investing heavily in MCTS.

    “You want to first make sure that this is good before you invest a lot of cycles doing mcts”

    Listen at 1:11:42

  18. Fifty thousand random games on a 9x9 board can train a reasonably good value function.

    “if you play like 50,000 games, you'll actually learn a pretty good value function as well”

    Listen at 1:13:46

  19. Go models can predict the winner despite being unable to predict the exact future board.

    “somehow we can predict who's going to win. And this captures a lot of possibilities here.”

    Listen at 1:21:30

  20. It remains unresolved whether tree structures can improve LLM reasoning.

    “the jury's still out as to whether this can ever work”

    Listen at 1:47:16

  21. Forward search and simulation may return as methods for improving AI reasoning.

    “the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback”

    Listen at 1:48:43

  22. High-dimensional control and language problems are less suited to Go-style discrete search heuristics.

    “most problems in much higher dimensional action spaces, or something that's combinatorially much bigger, like language, they don't seem as amenable to the kind of discrete action selection heuristics as well as kind of game evaluation type stuff that GO does”

    Listen at 1:50:27

  23. Eric Jangon Scaling lawsPositive1:53:33

    Scaling laws are most useful when the training recipe and dataset already work.

    “usually when you want scaling loss to work, you want to be in the regime where the recipe already works and the data sets are good”

    Listen at 1:53:33

  24. Creating a capability first generally requires more compute than reproducing it later.

    “the compute required to be the first to do something is always much larger than the compute it takes to catch up”

    Listen at 1:56:23

  25. Architecture choices have limited impact on current strong Go bot performance.

    “architecture choices don't matter that much”

    Listen at 1:59:16

  26. Desktop Blackwell GPUs can train Go bots using roughly half the GPU count of KataGo’s V100 setup.

    “Nvidia GPUs have indeed got faster. So whereas Katago was trained on V1 hundreds, you can train on half the number of desktop Blackwell GPUs and it still works.”

    Listen at 1:59:45

  27. Replay buffers should contain on-policy states plus off-policy recovery states.

    “your replay buffer really should have the states that your policy would visit, plus some distribution of states that you might drift to and then how to return back to your optimal states”

    Listen at 2:04:31

  28. Training on unreachable off-policy states wastes model capacity.

    “if the current model is looking at states that it would never reach, then it's kind of wasting capacity”

    Listen at 2:10:51

  29. Eric Jangon Soft labelsPositive2:18:55

    Soft labels contain more information than one-hot labels.

    “if you have access to the soft targets, the entropy of this distribution is far, far higher than the one hot”

    Listen at 2:18:55

  30. AlphaGo avoids starting reinforcement learning from zero success and solves exploration through improved labels.

    “you never have to initialize at a 0% success rate and solve the exploration problem of how to get a non zero success rate”

    Listen at 2:20:32

  31. Eric Jangon AlphaGo trainingPositive2:21:07

    AlphaGo’s supervised learning on improved labels is stable during training.

    “the training is very stable”

    Listen at 2:21:07

  32. Eric Jangon AI coding modelsPositive2:23:23

    Current AI coding models perform hyperparameter optimization effectively.

    “the models can do a very good job of doing hyperparameter optimization”

    Listen at 2:23:23

  33. Current publicly accessible closed models poorly select the next experiment within a research track.

    “current closed models that we can access, the public can access today, they don't seem to be that great at selecting what the next experiment should be in a given track”

    Listen at 2:25:28

  34. Eric Jangon AI compute scalingPositive2:31:23

    Over the long term, compute is likely the most important determinant of AI performance.

    “in the fullness of time, compute kind of is the single most important determinant on how things work”

    Listen at 2:31:23

  35. Experience developing game-playing AI may transfer positively to building language models.

    “there was some positive transfer from their time working on games and like Atari and Go and Starcraft that now helps them make good lms”

    Listen at 2:34:22

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Listen to the full episode and explore every guest, topic, and moment on PodLume.