← LLM inference latency
What podcasts say about LLM inference latency
Every statement, with the speaker, the exact quote and the moment it was said.
What experts have said about LLM inference latency
1 statement · 1 negative
Inference latency has a lower bound set by reading all model parameters from memory.
“there is a lower bound on latency, which is simply. Simply I need to read all of my total parameters from memory into the chips”
Open the episode · Reiner Pope – The math behind how LLMs are trained and servedListen at 10:28
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.