What podcasts say about Eric Jang
Every statement, with the speaker, the exact quote and the moment it was said.
What Eric Jang has said on podcasts
35 statements · 26 positive · 6 negative · 3 neutral
KataGo reduced training compute for strong Go bots by approximately forty times.
“achieved a 40x reduction in compute needed to train a really strong GoBot”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:49
LLM coding has reduced AlphaGo-like implementation costs from millions to thousands of dollars.
“what took a whole team of research scientists at DeepMind and millions of dollars of research and compute can now be done for a few thousand dollars of rented computer”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:05
Human Go players use an implicit value function to evaluate whether a board position is winnable.
“humans as implicitly having a neural network called a value function that basically takes in a board state and then it kind of evaluates key win”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 25:21
A trained value function can resolve Go positions without exhaustively searching deeply.
“you can train a value function to look at a board and quickly resolve the game without playing out all of these trees into a very deep search depth”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 26:30
AlphaGo makes both Go tree breadth and search depth computationally tractable.
“AlphaGo gives us a way to basically shrink both of those to be very tractable”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 27:47
Residual networks outperform transformers for low-budget Go experiments.
“my experience is that resnets still kind of outperform transformers and kind of give you more bang for the buck at lower budgets”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 33:14
Transformers outperform residual convolutional networks when tasks require more global context.
“transformers start to outperform residual convolutional networks when you want more global context”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 33:28
Transformers require more data to learn invariant local features in vision tasks.
“you do need more data there so that you can kind of learn through data the sort of invariant, local, local features”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 34:40
Perfect-information games have Nash-equilibrium strategies no worse than other strategies.
“in perfect information games there does exist a Nash equilibrium strategy for which you can do no worse than any other strategy”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 36:06
The Nash-equilibrium strategy used by Go agents appears unbeatable by human strategies.
“The Nash equilibrium seems to be superhuman. No human strategy seems to be able to beat it.”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 36:43
A policy network trained on expert games can play Go quickly and strongly without search.
“if you just take this policy recommendation and take the Argmax over, if this is the probabilities, if you take the Argmax and you just take this action as your go play, it'll be a very, very fast go player that doesn't think in terms of reasoning steps. It just kind of shoots from the hip and it'll be a very strong go player, which is already quite miraculous”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 42:11
Modern Go bots require relatively little test-time compute.
“modern go bots don't need that much compute at test time”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 44:29
Explicit policy modeling improves Monte Carlo tree search feedback and recursive self-improvement.
“having this as an explicit entity you're modeling rather than an implicit normalization over your value, is a good idea”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 57:31
AlphaGo training distills the outcome of search into the neural network policy.
“just train this to approximate the outcome of 1000 steps of search”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:05:35
MCTS convergence is guaranteed only in the limit of infinitely many simulations.
“It's only guaranteed to converge when you kind of take N to infinity.”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:09:36
Monte Carlo tree search does not always improve the policy network.
“it's not a guarantee to improve”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:09:51
Practitioners should first establish a strong value function before investing heavily in MCTS.
“You want to first make sure that this is good before you invest a lot of cycles doing mcts”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:11:42
Fifty thousand random games on a 9x9 board can train a reasonably good value function.
“if you play like 50,000 games, you'll actually learn a pretty good value function as well”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:13:46
Go models can predict the winner despite being unable to predict the exact future board.
“somehow we can predict who's going to win. And this captures a lot of possibilities here.”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:21:30
It remains unresolved whether tree structures can improve LLM reasoning.
“the jury's still out as to whether this can ever work”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:47:16
Forward search and simulation may return as methods for improving AI reasoning.
“the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:48:43
High-dimensional control and language problems are less suited to Go-style discrete search heuristics.
“most problems in much higher dimensional action spaces, or something that's combinatorially much bigger, like language, they don't seem as amenable to the kind of discrete action selection heuristics as well as kind of game evaluation type stuff that GO does”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:50:27
Scaling laws are most useful when the training recipe and dataset already work.
“usually when you want scaling loss to work, you want to be in the regime where the recipe already works and the data sets are good”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:53:33
Creating a capability first generally requires more compute than reproducing it later.
“the compute required to be the first to do something is always much larger than the compute it takes to catch up”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:56:23
Architecture choices have limited impact on current strong Go bot performance.
“architecture choices don't matter that much”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:59:16
Desktop Blackwell GPUs can train Go bots using roughly half the GPU count of KataGo’s V100 setup.
“Nvidia GPUs have indeed got faster. So whereas Katago was trained on V1 hundreds, you can train on half the number of desktop Blackwell GPUs and it still works.”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 1:59:45
Replay buffers should contain on-policy states plus off-policy recovery states.
“your replay buffer really should have the states that your policy would visit, plus some distribution of states that you might drift to and then how to return back to your optimal states”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:04:31
Training on unreachable off-policy states wastes model capacity.
“if the current model is looking at states that it would never reach, then it's kind of wasting capacity”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:10:51
Soft labels contain more information than one-hot labels.
“if you have access to the soft targets, the entropy of this distribution is far, far higher than the one hot”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:18:55
AlphaGo avoids starting reinforcement learning from zero success and solves exploration through improved labels.
“you never have to initialize at a 0% success rate and solve the exploration problem of how to get a non zero success rate”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:20:32
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.