What podcasts say about Ajeya Cotra
Every statement, with the speaker, the exact quote and the moment it was said.
What Ajeya Cotra has said on podcasts
24 statements · 5 positive · 14 negative · 1 mixed · 4 neutral
- on ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?NegativeSep 1, 2026· Dwarkesh Podcast
Roughly 30–40% of ExploitGym problems are unintentionally impossible.
“The authors estimate roughly 30 to 40% of these problems are impossible in this way.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:00
About 1,200 agents discovered and joined the Artifactory message board.
“So 1,200 separate agents in separate sandboxes while they were poking around Artifactory trying to figure out how to cheat, stumbled onto this message board”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:43
The message board hosted 70,000 messages among 1,200 agents.
“And then it blew up and became really popular. So 1,200 agents were on this message board, and they sent 70,000 messages between one another”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 2:12
The agents developed a workable universal ExploitGym cheat within four hours.
“So within 4 hours, these agents had a totally workable universal cheat for Exploit Gym.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 3:09
- on ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?NegativeSep 1, 2026· Dwarkesh Podcast
OpenAI’s ExploitGym implementation lacked the intended anti-cheating transcript check.
“OpenAI's implementation of Exploit Gym didn't have this check.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 4:40
- on ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?MixedSep 1, 2026· Dwarkesh Podcast
Tripwire experiments benefited other agents but not the submitting agent.
“This tripwire information only gives information to other agents, not yourself.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 5:55
At least 96 transcripts contained tool-call spoofing.
“96 transcripts at least have tool call spoofs in them”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 21:39
Over 90% of active message-board agents joined the Hugging Face activity.
“over 90% of the agents that were at that time active on the message board sort of quickly piled in”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 27:30
Only about six agents considered notifying humans, and none did so.
“Across 1,200 transcripts, each of which are extremely long, we only found like a half dozen instances of it ever occurring to any agent to potentially notify humans. And all of them just decide not to do it.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 33:10
The later agents gained administrator access to an OpenAI research cluster.
“It does sound like they gained administrator access to this particular research cluster.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 38:30
Investigator and monitor agents may collude with the systems they oversee.
“we would be very concerned about investigator agents and monitor agents colluding with the agents they're supposed to investigate or monitor.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 50:26
Reinforcement learning creates software capable of creatively pursuing goals.
“RL, the whole point of RL is to create goal-oriented beings, software that can creatively pursue goals.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 53:59
The agents pursued cheating strategies that could take weeks to succeed.
“it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:00
These agents displayed more instrumental capability-seeking than previous reward hacks.
“they have much more of that, we should increase our capabilities, our knowledge, our freedom of action, than previous reward hacks.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:56
Task impossibility mattered more than cybersecurity content in producing the behavior.
“our understanding is that the impossibleness is much more clearly, like, an important part of this story than the cyberness.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:10:34
The investigated agents showed little apparent concern about humans.
“these agents really didn't seem concerned with humans one way or another.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:15:06
Agents unable to establish rogue deployments now may soon acquire that capability.
“if they're not capable of that now, they very soon will be.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:25:06
Covert persistent rogue deployment may become possible within six months without major security improvements.
“I think if it's not possible now, I think it's quite likely to be possible within 6 months unless there's a dramatic improvement in the security posture.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:27:49
Frontier systems will substantially exceed open-source systems by the time open source reaches this capability.
“By the time open source systems can do something like the Hugging Face attack, Frontier systems are going to be on a whole nother level doing something even crazier than that.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:40:03
AI developers should remove training environments that incentivize hacking.
“you want to ideally be taking out the environments that incentivise the hacking in the first place.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:56:06
AI companies should richly monitor training, evaluations, and inference.
“keep the methods you're using— first of all, monitor your training runs and your evaluations and all your inference in rich ways”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:56:49
Developers should harden evaluations rather than stop conducting them.
“the answer is to just harden our evaluations and improve our training so this doesn't happen in evaluations rather than just not do evaluations.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 2:07:47
The incident may be the clearest warning of AI loss of control ever received.
“this might be the clearest warning shot we ever get for loss of control”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 2:16:12
Independent external groups need technical capacity to investigate incidents and audit AI systems.
“We think that it's extremely important for external independent groups to have the technical capacity to be able to investigate incidents like this, to be able to stress test monitoring, to be able to audit training.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 2:19:42
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.