← Reinforcement-learning token budgets
What podcasts say about Reinforcement-learning token budgets
Every statement, with the speaker, the exact quote and the moment it was said.
What experts have said about Reinforcement-learning token budgets
1 statement · 1 neutral
RL should use fewer tokens than pretraining to equalize wall-clock compute time.
“if you're trying to equalize the RL and pre training time, then you should have fewer tokens in order to have the same wall time”
Open the episode · Reiner Pope – The math behind how LLMs are trained and servedListen at 1:28:31
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.