← Training Data

What podcasts say about Training Data

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about Training Data

6 statements · 4 positive · 1 negative · 1 neutral

  1. Removing problematic training data from AI models is straightforward.

    “The alignment, you know, like you said, the first models are trained on every junk post on Reddit, every X tweet. And of course, they have weird things in there. But, you know, ferreting those out of the training data is so straightforward.”

    Listen at 52:43

    Open the episode · Frontier Labs Want to Slow Down, OpenAI Delays Its 2026 IPO, Anthropic Flags 5 Bioweapon Cases | EP #291
  2. Fei-Fei LiNegativeAug 10, 2026· Huberman Lab Essentials

    Limited training data was a major cause of slow AI progress.

    “So we conjectured that the lack of data was a huge part of the reason, that's the lack of progress in AI.”

    Listen at 10:52

    Open the episode · Using AI to Increase Your Intelligence & Enrich Humanity | Dr. Fei-Fei Li
  3. AI training data is generally commoditized rather than strategically proprietary.

    “My sense of data is that it's largely a commodity”

    Listen at 1:08:00

    Open the episode · Google's AI Brain Drain, SpaceX's Huge Quarter, Airtable's 90% Collapse, US Data Fuels China AI
  4. Ratings, critiques, curation, and examples each provide useful training data for AI models.

    “there's all these different data shapes that are helpful in different ways for model training.”

    Listen at 25:56

    Open the episode · Why AI has no taste and how to fix it (w/ Thais Castello Branco) | E2319
  5. Jen-Hsun HuangPositiveMar 23, 2026· Lex Fridman Podcast

    AI training data will continue scaling, increasingly using synthetic data.

    “we're going to keep on scaling the amount of data that we have to train with. A lot of that data is probably going to be synthetic.”

    Listen at 29:59

    Open the episode · #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution
  6. Dario AmodeiPositiveFeb 13, 2026· Dwarkesh Podcast

    Broad RL data is intended to produce generalization rather than teach isolated skills.

    “the goal is very similar to what was done five or 10 years ago with pre training with we're trying to get a whole bunch of data not because we want to cover a specific document or a specific skill, but because we want to generalize”

    Listen at 12:17

    Open the episode · Dario Amodei — The highest-stakes financial model in history

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

Training Data: what podcasts say · PodLume