What podcasts say about Beren Millidge
Every statement, with the speaker, the exact quote and the moment it was said.
What Beren Millidge has said on podcasts
17 statements · 8 positive · 4 negative · 2 mixed · 3 neutral
Failure to generalize meta-learning and continual learning could block transformative AI.
“if it is just ridiculously hard to generalize meta-learning, plus we don't solve continual learning, it's just super hard and impossible”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:37
AI systems are already close to crossing human-level capability on relevant performance measures.
“we're already pretty close, in my opinion, to where we'll start crossing the human Elo score”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 6:46
Rapid recursive self-improvement depends on AI systems learning their own objectives reliably.
“how well can AIs generalize to learning their own objectives”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 13:39
AI training environments can set tasks substantially beyond human capabilities.
“environments can go quite a far way above what humans can do”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 31:04
Simulation-to-real training will dominate while AI sample efficiency remains low.
“sim-to-real has to be the dominant framework while sample efficiency is kind of low”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 39:31
Chinese AI companies gain an advantage by training on deployment data and distilled model behavior.
“the Chinese 100% do. And they definitely get this advantage”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 42:39
It remains unresolved whether learned AI taste generalizes to very long-horizon tasks.
“how well does that generalize to really long horizon things is I think the question, which I think is really unsolved at this point”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 52:18
Continual learning may progress from quarterly releases to hourly updates, effectively solving deployment learning.
“instead of every 3 months we release a model, now it's every week and then every day and then every hour, at which point we basically have obviously solved it”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 54:52
Distilling new information into a separately trained base model is technically easier than continual updating.
“It's easy to distill it into a different base model”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 57:06
Continually mid-training one base model eventually reaches an asymptote.
“if you just keep continually mid-training the same base forever, it asymptotes at some point”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 59:20
Current RL explores poorly, making progress unlikely without success within roughly 128 rollouts.
“RL is not very good at exploring right now. And so if the model can't get in 128 rollouts, it's very unlikely to get signal to progress”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:04:20
Pretraining corpora cannot provide undiscovered solutions to frontier mathematical problems.
“There's no hidden proof of the Millennium Prize problem sitting in Common Crawl”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:05:50
Mid-training and post-training data become more valuable as model scale increases.
“a lot of the mid-training and post-training data we have now actually gets better with scale”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:09:38
Much apparent RL progress actually comes from strong synthetic mid-training data.
“An awful lot of what we see as successes of RL actually comes from very, very good mid-training data”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:18:57
RL entropy collapse mainly results from exploiting simplistic verifiers rather than RL itself.
“the RL entropy collapse is basically due to exploitation of fairly simple verifiers”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:28:09
Fully general AI remote workers may arrive in roughly three years.
“for the full generality, maybe 3 years”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:29:31
AI may reach superhuman performance in lab-focused domains within roughly five years.
“I kind of agree in the 5-year range, at least for the stuff that labs are focusing on”
Open the episode · AI researchers debate how close we are to recursive self-improvementListen at 1:36:23
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.