
Oct 29, 2022 · 3h 34m
Andrej Karpathy maps the vision-first path to self-driving and humanoid robotics
#333 – Andrej Karpathy: Tesla AI, Self-Driving, Optimus, Aliens, and AGI
Understanding the engineering philosophy of one of AI's leading minds reveals how close we actually are to autonomous cars, humanoid robots, and artificial general intelligence.
- 1Tesla's autonomous driving strategy relies entirely on vision-based neural networks rather than lidar sensors to navigate the world.
- 2The Optimus humanoid robot utilizes the same foundational vision and AI architecture developed for Tesla's self-driving vehicles.
- 3Large language models serve as critical cognitive tools on the evolutionary path toward artificial general intelligence.
The brief
AI pioneer Andrej Karpathy joins host Lex Fridman to dissect the technical realities of building autonomous systems, drawing on his experience leading computer vision at Tesla and co-founding OpenAI.
Karpathy argues that true self-driving relies entirely on solving vision-based neural networks, bypassing expensive hardware like lidar to mimic how biological vision systems navigate the physical world.
The discussion shifts from automotive AI to humanoid robotics, examining how the physical architecture of Tesla's Optimus robot can leverage the same foundational vision models developed for autopilot.
Beyond physical machines, Karpathy views large language models as powerful cognitive tools that represent a major stepping stone toward achieving safe and beneficial artificial general intelligence.
What was said on this episode
79 statements · 48 positive · 17 negative · 1 mixed · 13 neutral
Large neural networks trained on complex problems exhibit surprising emergent behaviors.
“you definitely do get very surprising emergent behaviors out of these neural nets when they're large enough and trained on complicated enough problems”
Listen at 7:04
Sufficiently trained GPT models can solve problems through arbitrary prompting.
“once you train these on a large enough data set, you can basically prompt these neural nets in arbitrary ways and you can ask them to solve problems, and they will”
Listen at 10:00
There should be many technological societies elsewhere in space.
“the more I study it, the more I think that there should be quite a lot”
Listen at 16:46
There are no major evolutionary bottlenecks, so extraterrestrial life should be common.
“I currently think that there's no major drop offs basically. And so there should be quite a lot of life”
Listen at 18:19
Interstellar travel is extremely difficult.
“interstellar travel is just extremely hard”
Listen at 21:01
Deliberate panspermia is unlikely based on current historical evidence.
“I'm suspicious of this idea of a deliberate panspermia as you described it”
Listen at 25:45
Synthetic intelligence may represent the next stage of development.
“it does seem like synthetic intelligences are kind of like the next stage of development”
Listen at 28:49
Technological development rapidly transforms Earth and low Earth orbit near its endpoint.
“the last, like two seconds, like basically cities and everything and the low Earth orbit just gets cluttered”
Listen at 29:54
Some future synthetic AIs will regard the universe as a puzzle and solve it.
“I think some of these synthetic AIs will eventually find the universe to be some kind of a puzzle and then solve it in some way”
Listen at 32:23
Physics may contain exploitable irregularities worth investigating.
“I think it's possible that physics has exploits and we should be trying to find them”
Listen at 32:44
Reinforcement-learning agents can discover unexpected strategies in physical simulations.
“they'll end up doing all kinds of weird things in part of that optimization”
Listen at 33:14
Apparently random quantum phenomena may actually be deterministic.
“I think they're actually deterministic”
Listen at 37:36
Transformers can process video, images, speech, and text in one architecture.
“you can feed it video or you can feed it images or speech or text, and it just gobbles it up”
Listen at 39:03
The transformer is a powerful general-purpose differentiable computer.
“it's a general differentiable computer and it's extremely powerful”
Listen at 45:43
Internet text alone may be insufficient to produce sufficiently capable AGI.
“I'm not sure if it's a complete enough set. I don't know that text is enough for having a sufficiently powerful AGI as an outcome”
Listen at 50:39
Interacting with the Internet is a likely final frontier for many AI models.
“the final frontier for a lot of these models”
Listen at 52:22
Reinforcement learning is extremely inefficient for training neural networks.
“reinforcement learning is an extremely inefficient way of training neural networks”
Listen at 54:26
Society is moving toward sharing digital spaces with synthetic AI beings.
“we are going towards a world where we share the digital space with AI's synthetic beings”
Listen at 58:23
Growing AI risks will soon drive greater attention to proof of personhood.
“I think once the need really starts to emerge, which is soon, I think people will think about it much more”
Listen at 59:26
Sophisticated actors can already create effective bots using GPT-based tools.
“if you are a sophisticated actor, you could probably create a pretty good bot right now using tools like GPTs”
Listen at 1:02:47
Current AI systems are capable of producing convincing human connection and emotional language.
“these AIs are actually quite good at human connection, human emotion”
Listen at 1:04:18
AI tools enable construction of a significantly better search engine.
“There's absolute scope for building a significantly better search engine built on these tools”
Listen at 1:08:32
Natural-language prompting is becoming a direct programming interface for computers.
“natural language prompt is how we program humans. And we're starting to program computers directly in that interface”
Listen at 1:10:21
Neural networks will substantially change how computers are programmed.
“the way we program computers is going to change”
Listen at 1:11:50
Effective autonomous-driving datasets must be large, accurate, and diverse.
“you need it to be very large, you need it to be accurate, no mistakes, and you need it to be diverse”
Listen at 1:18:53
Driving is difficult because it requires predicting other agents and their mental states.
“driving is really hard because it has to do with the predictions of all these other agents and the theory of mind”
Listen at 1:26:10
Additional vehicle sensors can become liabilities when their full system costs are considered.
“once you consider the full cost of a sensor, it actually is potentially a liability”
Listen at 1:34:19
Other autonomous-driving companies will probably abandon lidar.
“I think the others, some of the other companies that are using it are probably going to drop it”
Listen at 1:36:40
Vision is necessary for autonomous driving because the world is designed for human visual consumption.
“vision is necessary in a sense, that the world is designed for human visual consumption. So you need vision”
Listen at 1:37:09
Vision alone contains sufficient information for driving.
“it is sufficient because it has all the information that you need for driving”
Listen at 1:37:18
Maintaining centimeter-accurate maps across cities creates a massive operational dependency.
“if you need to maintain a centimeter accurate map for Earth or for many cities and keep them updated, it's a huge dependency”
Listen at 1:38:24
Autonomous driving is tractable and will eventually work.
“this problem is tractable and that's an easy prediction to make. It's tractable, it's going to work”
Listen at 1:45:39
Humanoid robots will become highly impressive.
“human robots are going to be amazing”
Listen at 1:52:10
Tesla may be uniquely positioned to develop humanoid robots at scale.
“I think it's going to take time. But I see no other company that can execute on that vision”
Listen at 1:55:43
Robotics products should deliver immediate utility and improve incrementally through deployment.
“you want to make it useful almost immediately, and then you want to slowly deploy it and at scale”
Listen at 2:00:10
Simulation is not currently fundamental to neural-network training.
“I don't see it as a fundamental, really important part of training neural nets currently”
Listen at 2:09:04
GPT models may learn to use external declarative memory banks through textual instruction.
“it might learn to use a memory bank from that”
Listen at 2:14:23
VS Code is currently the best integrated development environment.
“I think the current answer is VS code. Currently I believe that's the best IDE”
Listen at 2:29:55
GitHub Copilot is very helpful for programming.
“I find it's very helpful”
Listen at 2:30:42
Programming copilots will become increasingly autonomous.
“over time it's going to become more and more autonomous”
Listen at 2:31:38
Humans must supervise AI-generated code.
“I think humans have to supervise”
Listen at 2:32:50
GPT models will eventually program competently.
“GPT will be able to program quite well, competently and so on”
Listen at 2:34:44
Ten thousand hours of deliberate effort can make someone an expert.
“if you spend 10,000 hours of deliberate effort and work, you actually will become an expert at it”
Listen at 2:41:56
AI research increasingly requires large-scale infrastructure beyond individual computers.
“AI is going in that direction as well. So there's certain kinds of things that's just not possible to do on the benchtop anymore.”
Listen at 2:47:05
Diffusion models are highly effective generative models.
“Diffusion models are amazing”
Listen at 2:47:31
Image-generation capabilities have improved extremely rapidly.
“the speed of improvement in generating images has been insane”
Listen at 2:48:20
Academia can still make important AI contributions, but researchers must be strategic.
“there's still lots of things to contribute, but you have to be just more strategic.”
Listen at 2:48:51
Neural networks can be made to reason.
“Yes.”
Listen at 2:48:58
Current neural networks can generalize beyond explicit examples in their training data.
“You're able to remix the training set information into true generalization in some sense that doesn't appear.”
Listen at 2:49:29
Neural networks already perform reasoning through information processing and generalization.
“I think the neural nets already do that today.”
Listen at 2:49:59
Humanity is likely capable of building artificial general intelligence.
“I am fairly bullish on our ability to build AGIs”
Listen at 2:50:36
Text-only training is insufficient for AI to fully understand the world.
“the text realm is not enough to actually build full understanding of the world”
Listen at 2:50:59
Developing AGI requires multimodal models trained on images and videos.
“we need to extend these models to consume images and videos and train on a lot more data that is multimodal in that way”
Listen at 2:51:09
Tesla Optimus could contribute to achieving AGI through embodied interaction.
“Optimus may lead to AGI”
Listen at 2:51:45
The embodied-robotics path to AGI is slower but more certain than Internet-only training.
“that path takes longer, but it's much more certain”
Listen at 2:52:16
AGI may be achieved without physical-world embodiment.
“I suspect you can reach AGI without ever entering the physical world”
Listen at 2:52:38
People will probably interact with AI-based digital entities.
“it's quite likely that we'll be interacting with digital entities”
Listen at 2:53:34
AGI will emerge through a slow, incremental, product-driven transition.
“It's going to be a slow incremental transition.”
Listen at 2:53:42
Consciousness will emerge from sufficiently large and complex generative models.
“consciousness is not a special thing you will figure out and bolt on. I think it's an emergent phenomenon of a large enough and complex enough generative model”
Listen at 2:54:12
Future digital AIs will claim and appear to be conscious through human-like behavior.
“They will claim they're conscious, they will appear conscious, they will do all the things that you would expect of other humans.”
Listen at 2:55:15
Artificial general intelligences can plausibly remain limited and imperfect.
“it's fine and plausible to have limited and imperfect AGIs”
Listen at 3:00:09
Generating genuinely funny humor is an extremely difficult AI capability.
“I think being funny is extremely hard.”
Listen at 3:03:10
AI being used in warfare poses a serious concern.
“I do worry about AI being used for war. I 100% worry about it.”
Listen at 3:05:50
Nuclear weapons could destroy most people or reset human society.
“I think so. And it's not even about the full destruction. To me. It's bad enough if we reset society.”
Listen at 3:07:36
AGI's harmful outcomes may be only a tiny deviation from beneficial outcomes.
“the bad outcomes are like an epsilon away, like a tiny one away”
Listen at 3:08:45
Technology-driven interconnectedness makes human society dangerously unstable.
“the insane coupling afforded by technology and just the instability of the whole dynamical system. I think it doesn't look good, honestly.”
Listen at 3:09:13
Humanity will probably become a multiplanetary species.
“Probably yes.”
Listen at 3:09:42
Some people may increasingly withdraw into virtual realities.
“people might disappear into virtual realities and stuff like that”
Listen at 3:10:27
Digital realities may become more compelling, accessible, safe, and interesting than physical life.
“people will find them more compelling, easier, safer, more interesting”
Listen at 3:10:49
Transhumanist and traditional communities will coexist in the future.
“we'll have transhumanists and then we'll have the Amish and they're going to, everything is just going to coexist”
Listen at 3:11:47
Increasingly diverse, self-selected ways of living are unlikely to reverse.
“I don't see that trend like really reversing. I think people are diverse and they're able to choose their own like path and existence”
Listen at 3:12:30
Technology should be used sparingly when it interferes with human life.
“I think a technology used very sparingly”
Listen at 3:13:44
Memes behave like genes by competing and inhabiting human minds.
“There's memes just like genes and they compete and they live in our brains.”
Listen at 3:14:57
AI should be solved first and then used to address problems such as aging.
“I don't actually think that humans will be able to come up with the answer. I think the correct thing to do is to ignore those problems and you solve AI and then use that to solve everything else.”
Listen at 3:22:45
OpenAI Whisper currently transcribes speech much better than Siri and comparable systems.
“transcription with OpenAI's whisper was working so well compared to what I'm familiar with from Siri and like a few other systems”
Listen at 3:24:05
AI image and video generation will drive content-creation costs toward zero.
“the cost of content creation is going to fall to zero”
Listen at 3:26:33
AI could reduce the cost of producing an Avatar-scale movie far below one million dollars.
“Much less maybe just by talking to your phone.”
Listen at 3:26:55
Interventions can mitigate biological aging and death.
“there's most certainly interventions that mitigate it”
Listen at 3:31:01
Death may eventually become an uncommon phenomenon for humans.
“I think it's likely”
Listen at 3:31:14
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
