← LLM prefill and decode

What podcasts say about LLM prefill and decode

Every statement, with the speaker, the exact quote and the moment it was said.

What experts have said about LLM prefill and decode

1 statement · 1 neutral

  1. Reiner PopeNeutralApr 29, 2026· Dwarkesh Podcast

    Prefill is compute-limited, whereas decode is memory-bandwidth-limited.

    “prefill is compute limited and decode is memory bandwidth limited”

    Listen at 1:45:11

    Open the episode · Reiner Pope – The math behind how LLMs are trained and served

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

LLM prefill and decode: what podcasts say · PodLume