Dwarkesh Podcast
Dwarkesh Podcast

Dec 23, 2025 · 12 min

Post-training and reinforcement learning emerge as the new AI frontier

An audio version of my blog post, Thoughts on AI progress (Dec 2025)

As traditional scaling methods encounter practical limits, the AI industry must pivot to new training paradigms to sustain rapid cognitive progress.

3 key takeaways
  1. 1Post-training techniques like reinforcement learning are becoming the primary drivers of model intelligence over raw data scaling.
  2. 2Achieving artificial general intelligence may require models to master multi-step planning and autonomous self-correction.
  3. 3The future of AI development is shifting toward agentic systems that operate like specialized, cooperative software.

The brief

The narrative surrounding artificial intelligence is shifting from raw scaling power to the complex engineering of post-training, where reinforcement learning and specialized fine-tuning are redefining the limits of large language models.

While early breakthroughs relied heavily on massive pre-training datasets, the next frontier of progress lies in teaching models how to reason, self-correct, and execute multi-step planning in dynamic environments.

This shift raises critical questions about whether current deep learning architectures can achieve artificial general intelligence, or if entirely new paradigms in computational neuroscience and system design are required.

Ultimately, the path to advanced AI may look less like a single massive brain and more like highly specialized software systems, where models act as cooperative agents solving targeted, highly complex tasks.

What was said on this episode

17 statements · 5 positive · 8 negative · 3 mixed · 1 neutral

  1. Training on verifiable outcomes is doomed if human-like learners are near.

    “If we're actually close to a human-like learner, then this whole approach of training on verifiable outcomes is doomed.”

    Listen at 0:07

  2. Self-directed on-the-job learning would make extensive skill pre-training pointless; otherwise AGI is not imminent.

    “either these models will soon learn on the job in a self-directed way, which will make all this pre-baking pointless, or they won't, which means that AGI is not imminent.”

    Listen at 0:31

  3. AIs currently lack a robust, efficient way to acquire company-specific job skills.

    “there just isn't currently a robust, efficient way for AIs to pick up these skills.”

    Listen at 3:06

  4. Useful automation requires AI that learns from feedback or experience and generalizes like humans.

    “What you actually need is an AI that can learn from semantic feedback or from self-directed experience and then generalize the way a human does.”

    Listen at 4:14

  5. Predefined skills alone cannot automate even one complete job.

    “It is not possible to automate even a single job by just baking in a predefined set of skills, let alone all the jobs.”

    Listen at 4:36

  6. Brain-like artificial intelligences will emerge within one or two decades.

    “I expect actual brain-like intelligences within the next decade or two”

    Listen at 4:59

  7. Human-like AI models would diffuse rapidly across firms.

    “If these models actually were like humans on a server, they'd diffuse incredibly quickly.”

    Listen at 5:25

  8. AI labor will diffuse into firms more easily than human labor is hired.

    “I expect it's going to be much easier to diffuse AI labor into firms than it is to hire a person”

    Listen at 6:03

  9. AGI-level models would command trillions of dollars in annual token spending.

    “If the capabilities were actually at AGI level, people would be willing to spend trillions of dollars a year buying tokens that these models produce.”

    Listen at 6:11

  10. Current AI models are far less capable than human knowledge workers.

    “the reason that labs are orders of magnitude off this figure right now is that the models are nowhere near as capable as human knowledge workers.”

    Listen at 6:25

  11. By 2030, continual-learning progress and hundreds of billions in revenue will coexist with incomplete automation.

    “I expect that by 2030 that labs will have made significant progress on my hobby horse of continual learning and the models will be earning hundreds of billions of dollars in revenue a year, but they won't have automated all knowledge work”

    Listen at 7:52

  12. AI models are becoming impressive quickly but useful more slowly.

    “Models keep getting more impressive at the rate of the short timelines people predict, but more useful at the rate that the long timelines people predict.”

    Listen at 8:13

  13. No well-fitting public scaling trend is known for reinforcement learning from verifiable reward.

    “for which we have no well-fit publicly known trend.”

    Listen at 8:51

  14. Human-level AI on-the-job learning may require another five to ten years.

    “human-level, on-the-job learning may take another five to 10 years to iron out.”

    Listen at 10:55

  15. The first model achieving continual learning will not produce runaway gains.

    “I don't expect some kind of runaway gains from the first model that cracks continual learning”

    Listen at 11:02

  16. Competition among major model companies will remain intense.

    “the competition will stay pretty fierce between all of these model companies”

    Listen at 11:34

  17. Chat engagement and synthetic-data flywheels have barely reduced competition among model companies.

    “all of these previous supposed flywheels, whether that's user engagement on chat or synthetic data or whatever, have done very little to diminish the greater and greater competition between model companies.”

    Listen at 11:39

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Books & mentions

Listen to the full episode and explore every guest, topic, and moment on PodLume.

Post-training and reinforcement learning emerge as the new AI frontier · PodLume