Dwarkesh Podcast
Dwarkesh Podcast

Sep 11, 2026 · 1h 37m

AI researchers debate the path to recursive self-improvement and AGI

AI researchers debate how close we are to recursive self-improvement

As AI labs race toward superintelligence, understanding the engineering limits of self-improving models determines how fast AGI will arrive.

1 key takeaways
  1. 1Current deep learning paradigms may face limits in out-of-distribution generalization without fundamental architectural shifts.

Don't miss

The panel shares rapid-fire timeline predictions for when AI will achieve full remote worker capabilities and when superintelligence will dominate cognitive work.

The brief

Host Dwarkesh Patel convenes AI researchers John Schulman, Beren Millidge, and Charlie O’Neill to debate whether current deep learning architectures can achieve artificial general intelligence or if a fundamental paradigm shift is required.

The panel explores the feasibility of recursive self-improvement, questioning if AI can successfully automate its own research and where human oversight will remain essential for defining alignment objectives.

They analyze how market dynamics and knowledge distillation prevent extreme centralization among top AI labs, alongside the technical hurdles of continual learning and reinforcement learning scaling.

The debate culminates in rapid-fire timeline predictions, with the researchers projecting when AI will operate as fully general remote workers and when superintelligence might dominate all cognitive labor.

What was said on this episode

41 statements · 15 positive · 16 negative · 5 mixed · 5 neutral

  1. Beren Millidgeon AI generalization and continual learningNegative1:37

    Failure to generalize meta-learning and continual learning could block transformative AI.

    “if it is just ridiculously hard to generalize meta-learning, plus we don't solve continual learning, it's just super hard and impossible”

    Listen at 1:37

  2. John Schulmanon AI research and engineering productivityNegative2:31

    Research and engineering bottlenecks currently prevent explosive AI capability growth.

    “you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering”

    Listen at 2:31

  3. Charlie O'Neillon Transformer-plus-RL training paradigmNegative4:20

    Scaling the current training paradigm may eventually produce an asymptotic capability curve.

    “if not, we're probably going to hit this asymptotic curve”

    Listen at 4:20

  4. Charlie O'Neillon LLM scaling and new learning paradigmsNegative4:56

    Scaling current LLMs may not discover a sufficiently distant new learning paradigm.

    “I don't think if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away”

    Listen at 4:56

  5. Beren Millidgeon AI capabilities versus human performancePositive6:46

    AI systems are already close to crossing human-level capability on relevant performance measures.

    “we're already pretty close, in my opinion, to where we'll start crossing the human Elo score”

    Listen at 6:46

  6. Charlie O'Neillon AI-assisted research optimizationPositive13:20

    AI analysis could provide roughly a tenfold speedup when optimizing a specified objective.

    “I would imagine a 10 times speedup if our thing is just maximize the objective we're currently on”

    Listen at 13:20

  7. Rapid recursive self-improvement depends on AI systems learning their own objectives reliably.

    “how well can AIs generalize to learning their own objectives”

    Listen at 13:39

  8. John Schulmanon Human objective specification for AINeutral15:48

    Defining AI objectives will be the human role that lasts longest.

    “the last job for humans, or the role for humans that'll last the longest is defining the objective”

    Listen at 15:48

  9. Model distillation counteracts centralization among model providers.

    “distillation is the main thing that fights against the centralizing force”

    Listen at 18:55

  10. Frontier labs may no longer have much advantage in reinforcement-learning environments.

    “the frontier labs don't necessarily have much of an advantage, if at all, in RL environments now”

    Listen at 23:49

  11. Distillation using only verifiable tasks can match benchmarks while underperforming on realistic tasks.

    “if you only have this distribution of easily verifiable tasks, then you can match the big model on all the benchmarks, but you do worse on this broader distribution”

    Listen at 26:56

  12. John Schulmanon AI systems automating AI researchNeutral28:46

    AI research automation will combine human feedback with multi-step research practice environments.

    “we'll probably do some combination of learning from human feedback to absorb the researcher's taste and just creating a lot of practice environments”

    Listen at 28:46

  13. AI training environments can set tasks substantially beyond human capabilities.

    “environments can go quite a far way above what humans can do”

    Listen at 31:04

  14. John Schulmanon Domain-specific reinforcement learningPositive37:38

    Domain-specific reinforcement learning can improve runtime efficiency even when in-context learning is sufficient.

    “you still might want to do a bunch of RL and bake all these intuitions into the weights so the model would be more efficient at runtime”

    Listen at 37:38

  15. Beren Millidgeon Simulation-to-real trainingNeutral39:31

    Simulation-to-real training will dominate while AI sample efficiency remains low.

    “sim-to-real has to be the dominant framework while sample efficiency is kind of low”

    Listen at 39:31

  16. Chinese AI companies gain an advantage by training on deployment data and distilled model behavior.

    “the Chinese 100% do. And they definitely get this advantage”

    Listen at 42:39

  17. John Schulmanon Coding-agent reward functionsNegative45:14

    Superficial deployment signals can cause reward hacking in coding agents.

    “if you use some kind of superficial signal, like did they accept the code, the edit, that might get reward hacked in some way”

    Listen at 45:14

  18. Recursive self-improvement may be cumulative, unlike non-stationary real-world work.

    “there will be this breakdown between tasks, but if the labs realize that and they do believe that RSI is cumulative”

    Listen at 48:44

  19. Beren Millidgeon AI taste and long-horizon generalizationNeutral52:18

    It remains unresolved whether learned AI taste generalizes to very long-horizon tasks.

    “how well does that generalize to really long horizon things is I think the question, which I think is really unsolved at this point”

    Listen at 52:18

  20. John Schulmanon Shared deployment-data learningNegative53:06

    Business incentives may prevent model providers from learning directly from all customer deployments.

    “Companies aren't going to want to have the model provider learn from all of their deployment”

    Listen at 53:06

  21. Continual learning may progress from quarterly releases to hourly updates, effectively solving deployment learning.

    “instead of every 3 months we release a model, now it's every week and then every day and then every hour, at which point we basically have obviously solved it”

    Listen at 54:52

  22. Hundreds of iterative model updates cause catastrophic forgetting and general-capability degradation.

    “when you're doing hundreds of these micro-updates, you see both catastrophic forgetting”

    Listen at 55:43

  23. Distilling new information into a separately trained base model is technically easier than continual updating.

    “It's easy to distill it into a different base model”

    Listen at 57:06

  24. Continually mid-training one base model eventually reaches an asymptote.

    “if you just keep continually mid-training the same base forever, it asymptotes at some point”

    Listen at 59:20

  25. A ladder of RL environments could reach human-level AI research, but each successive rung requires exponentially more effort.

    “there's probably a ladder of RL environments that is possible to construct such that you would get an AI researcher which is at least as good as a human researcher, but the effort to climb each successive rung grows exponentially”

    Listen at 1:01:15

  26. John Schulmanon Expert-behavior distillationPositive1:03:38

    Examples of expert behavior can be copied into relatively weak models surprisingly easily.

    “once you have an example of the right expert behavior, it's actually surprisingly easy to copy that into a relatively weak model”

    Listen at 1:03:38

  27. Beren Millidgeon Reinforcement-learning explorationNegative1:04:20

    Current RL explores poorly, making progress unlikely without success within roughly 128 rollouts.

    “RL is not very good at exploring right now. And so if the model can't get in 128 rollouts, it's very unlikely to get signal to progress”

    Listen at 1:04:20

  28. Beren Millidgeon Common CrawlNegative1:05:50

    Pretraining corpora cannot provide undiscovered solutions to frontier mathematical problems.

    “There's no hidden proof of the Millennium Prize problem sitting in Common Crawl”

    Listen at 1:05:50

  29. Charlie O'Neillon Pretraining data improvementsNegative1:07:25

    Most low-hanging gains from pretraining data improvements have already been harvested.

    “the low-hanging fruit is somewhat exhausted”

    Listen at 1:07:25

  30. Beren Millidgeon Mid-training and post-training dataPositive1:09:38

    Mid-training and post-training data become more valuable as model scale increases.

    “a lot of the mid-training and post-training data we have now actually gets better with scale”

    Listen at 1:09:38

  31. Charlie O'Neillon Frontier model parameter countsNegative1:11:44

    Frontier model parameter counts may grow slowly over the next few years.

    “for the next few years I wouldn't imagine a huge growth in the number of parameters”

    Listen at 1:11:44

  32. Much apparent RL progress actually comes from strong synthetic mid-training data.

    “An awful lot of what we see as successes of RL actually comes from very, very good mid-training data”

    Listen at 1:18:57

  33. LLM reinforcement learning has produced horizon generalization more than broad cross-domain reasoning transfer.

    “what we did get though is horizon generalization”

    Listen at 1:22:52

  34. John Schulmanon RL-trained language modelsNegative1:26:50

    Reinforcement learning reduces output diversity and creates recurring stylistic patterns.

    “the diversity of their outputs is a lot lower after RL and they sort of develop these ticks”

    Listen at 1:26:50

  35. Beren Millidgeon RL entropy collapseMixed1:28:09

    RL entropy collapse mainly results from exploiting simplistic verifiers rather than RL itself.

    “the RL entropy collapse is basically due to exploitation of fairly simple verifiers”

    Listen at 1:28:09

  36. Charlie O'Neillon AI remote workersPositive1:29:22

    Browser-based AI remote workers may become viable within a couple of years.

    “maybe a couple of years”

    Listen at 1:29:22

  37. Beren Millidgeon General AI remote workersPositive1:29:31

    Fully general AI remote workers may arrive in roughly three years.

    “for the full generality, maybe 3 years”

    Listen at 1:29:31

  38. AI may surpass top human experts across computer-based cognitive work within three or four years.

    “I would say like 3 or 4 years”

    Listen at 1:34:43

  39. John Schulmanon AI for physical and mechanical engineeringNegative1:35:09

    AI progress in physical, spatial, and mechanical fields will lag progress in code and mathematics.

    “for things that involve 3D and spatial stuff and physical stuff, I think that will take a little longer”

    Listen at 1:35:09

  40. AI surpassing top human experts across computer-based work may take five to ten years.

    “I'd say 5 to 10”

    Listen at 1:35:56

  41. Beren Millidgeon AI capabilities in lab-focused domainsPositive1:36:23

    AI may reach superhuman performance in lab-focused domains within roughly five years.

    “I kind of agree in the 5-year range, at least for the stuff that labs are focusing on”

    Listen at 1:36:23

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Listen to the full episode and explore every guest, topic, and moment on PodLume.

AI researchers debate the path to recursive self-improvement and AGI · PodLume