
Jun 26, 2026 · 20 min
AI labs bet on reinforcement learning to solve long-horizon tasks
The next big breakthrough will be AIs learning on the job
The core research bet of major AI labs relies on reinforcement learning, but technical bottlenecks in context length and generalization could stall the path to true autonomy.
- 1AI labs are training agents across millions of verifiable environments to achieve artificial general intelligence.
- 2Serving models at longer context lengths than they were trained on leads to significant performance degradation.
- 3Short-horizon reinforcement learning training may fail to scale to complex, open-ended real-world endeavors.
The brief
Major AI labs are betting their futures on training agents across millions of verifiable tasks and reinforcement learning environments to unlock artificial general intelligence.
A critical technical bottleneck lies in context length, where training models at a short context length but serving them at a longer one causes severe performance degradation.
The ultimate test for reinforcement learning is generalization, specifically whether training on short-horizon tasks can scale to long-horizon, open-ended real-world challenges.
While AI agents may master specific white-collar tasks, it remains highly uncertain if they can generalize to highly complex, open-ended endeavors like building a business.
What was said on this episode
24 statements · 16 positive · 5 negative · 3 neutral
Sufficient RLVR training across diverse verifiable tasks will produce AGI.
“They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments, then we will have basically built AGI.”
Listen at 0:02
RL training is improving agents’ ability to solve ambitious, long-horizon problems.
“AI agents are able to solve more and more ambitious problems over longer and longer time spans.”
Listen at 1:17
Sufficiently strong in-context learning could eliminate the need for continual weight updates.
“if in context learning gets so good across longer and longer time horizons, then you don't need to distill back everything the model is learning on the job into the weights.”
Listen at 1:35
AI systems may achieve effectively infinite-feeling context windows within several years.
“with a couple more years of progress, we might have what feels like infinitely large context windows.”
Listen at 2:07
Verifiability alone is insufficient; domains must also support grindable training.
“it is not enough for a domain to be verifiable. It's it also has to be very grindable”
Listen at 3:09
Highly capable coding AIs will accelerate computer-use progress by building faithful application clones.
“Of course, once AIs get good enough at coding themselves to build these clones with extremely high fidelity, then I'm sure the computer use will make quicker progress than it is right now.”
Listen at 4:01
AI computer use may be solved soon.
“while computer use itself may soon be solved”
Listen at 4:20
Replayable training targets are necessary for substantial model progress in a domain.
“unless you can build a very replayable training target for a domain, the models will struggle to make much progress.”
Listen at 4:25
Sparse, idiosyncratic real-world data makes sample efficiency necessary for AI proficiency.
“because of the idiosyncratic and sparse nature of data in most domains in the world, you need sample efficiency in order to get proficient.”
Listen at 5:32
AI labs expect RLVR-trained agents to generalize beyond their training environments.
“The labs are betting that RL VR will generalize”
Listen at 6:09
Training on reproducible environments will create general agents that execute plans and learn rapidly.
“if you train on enough containerized, reproducible environments, you will develop a very general agent that can make it execute plans and learn rapidly from new information”
Listen at 6:13
Short-horizon RL training may fail to generalize to long-horizon performance.
“it seems like he's saying that short horizon RL training doesn't necessarily generalize to long horizon RL performance.”
Listen at 7:18
Inference consumes 30–50% of lab compute and currently does not improve models.
“Around 30 to 50% of a lab's compute goes to inference and that compute is currently not playing any productive role in helping improve the model.”
Listen at 7:50
Deployment reveals the most valuable information for improving AI models.
“it is only in deployment that the most valuable bits of information which your model could learn from are actually revealed.”
Listen at 8:01
Continual learning through indefinitely expanding KV caches is not scalable.
“AIs can't just keep building up a bigger and bigger KV cache as they learn from more and more users. That's just not scalable”
Listen at 8:45
Deployed online-learning models currently require repeated identical objectives across millions of users.
“all of the successfully shipped online learning models have had to learn the exact same thing across millions of users.”
Listen at 9:46
Model architecture is probably not the fundamental bottleneck for continual learning.
“It doesn't seem to me that architecture is fundamentally what is bottlenecking continual learning.”
Listen at 12:37
On-policy self-distillation can improve continual learning without requiring an outer-loop verifiable reward.
“This is better than RLVR for two reasons. One OPSD doesn't require us to have some outer loop verifiable reward.”
Listen at 13:28
On-policy self-distillation provides denser supervision than naive reinforcement learning.
“OPSD provides a much denser supervision signal than naive RO.”
Listen at 13:49
Reinforcement learning concentrates model updates on information relevant to achieving outcomes.
“RL is great at concentrating the update to only what is relevant to getting the outcome right.”
Listen at 14:32
Reality simulations could provide AI with vastly more training samples without additional wall-clock time.
“If the AI can build a good simulation of reality against which to rehearse new skills, or try alternative strategies and reinforce what actually works, then AIs could experience all orders of magnitude more simulated samples in the same wall clock time.”
Listen at 15:36
Successful AI dreaming would become a fourth major scaling axis.
“If it works, it would become a fourth axis of scaling alongside pre training, RL and inference time compute”
Listen at 16:43
Mature AI systems will improve mainly through experience accumulated after deployment.
“By this point, the main way that AIs get better is not from the training they have received before they are released to the public. Rather, it's from all this experience that they'll be accumulating from being broadly deployed in the economy”
Listen at 19:08
Future AIs will become smarter through learning from each user interaction worldwide.
“Every time that you interact with an AI, it'll be smarter, not only because it's been learning from your previous sessions, but also because it's been learning from all its interactions with all the other users in the world.”
Listen at 19:22
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
Books & mentions
Dwarkesh Podcast
Deep-dive interviews on the technical frontiers of artificial intelligence and policy