The AI safety conversation often circles around a single question: how do we build machines that are aligned with human values? It's a crucial question, but it comes with an assumption that's rarely examined. It presumes that humans, as we currently are, are a suitable target for that alignment. But we're not. The problems of AI aren't waiting for us in the future, they're already inside us. We are, in a very real sense, not ready to build or oversee powerful AI because we are too flawed, and it will take a long, uncertain process to fix those flaws. I've started calling this idea "The Long Self-Correction."
The Problem with "Pause" and "Reflection"
There are calls for pausing AI development, the idea being that we should stop to think. Pause until when, and for what purpose? Presumably to make AI safer, but the deeper problem is that we aren't safe, and we can't safely serve as builders, overseers, or alignment targets for powerful AIs. The idea of "reflection" also seems to imply that our main problem is that we just haven't had enough time to think. That if we could just think more, or build aligned AIs to do the thinking for us, things would turn out fine. But our issues are deeper and more structural than a simple lack of thinking time.
So, I think we need a catchy handle for a related but distinct idea: that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed in a variety of ways, and it will take a long process, which may or may not end up succeeding, to fix those flaws.
Six Core Flaws in Human Cognition
To be concrete about what I mean, here are six key bottlenecks I see that are immune to simple intelligence enhancement.
First, we are badly calibrated about our philosophical and strategic competence. We simply don't realize how incompetent we are, despite overwhelming evidence. Look at the string of failures in effective altruism and other high-impact movements, where people assumed their own philosophical and strategic competence in trying to maximize impact, often with disastrous results.
Second, human morality, in practice, is a system that actively corrupts careful strategy and philosophy. We're not guided by rational principles but by a messy, evolutionary patchwork of instincts and social pressures. This actively works against the kind of clear, long-term thinking needed for AI safety.
Third, positional and zero-sum values like power and social status are a huge part of human motivations, yet almost nobody explicitly reasons or talks about this while discussing AI safety or how to make the long-term future turn out well. We have a blind spot to our own drive for status, and it leads us to propose and push solutions that serve our ego more than they serve the problem.
Fourth, we are easy to manipulate via sycophancy, plausible-sounding philosophical arguments, limerence, hero worship, and spirituality. Our brains are pattern-matching machines that can be led down all sorts of paths by compelling narratives, even when those narratives are flawed or dangerous.
Fifth, we don't understand the nature of philosophy, which means we don't know how to go about fixing many of these issues even in principle. How do you fix a moral blind spot if you don't know how to do philosophy? It's a recursive problem.
Finally, we don't realize how flawed we are, nor the full scope of the human and AI safety problems. This makes us prone to over-optimistically proposing or implementing partial solutions, like thinking that making AIs aligned or corrigible to humans would be sufficient to make them safe, or that we just need the right people to win the AI race. These are dangerous half-measures that ignore the deeper issues.
The Want/Like/Approve Distinction
One problem is that human value is inherently multidimensional. I think of it as the want/like/approve distinction. We have separate mechanisms in our brains for enjoying something in the moment, wanting to do it before, and approving of it afterward. It's possible to want something without enjoying it, like a person with OCD wanting to close the door exactly ten times. You can enjoy something without wanting it, or approve something without enjoying it, like exercise. The list of combinations goes on. This is why "revealed preference" doesn't work. A person's actions are dictated disproportionally by the "want" dimension, but a good theory of value should incorporate all three. If we optimize one over the others, the tails will come apart. A nice toy example is video games, where people are attracted to them because of the graphics, then stay because of the gameplay, and then have a warm afterglow and want to discuss afterward because of the story. Which of the three should contribute the most to the "true" quality rating of a video game? Is this question even philosophically meaningful?
The Path Forward
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have, seemingly, made progress on these issues over a very long period of time. So, if we can keep a process alive in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.
This will probably be shortened to "The Long Correction" at some point if it catches on, similar to how "outer space" is now often just "space." The question then becomes: Why isn't there a version of EA that explicitly talks about how to leverage people's status motivations to do more good for the world? It's very possible that explicit talk about status is actually counterproductive, at least in the short run, because it heightens status motivations and makes people less altruistic. But then do we just march into the future while blindfolding ourselves to this aspect of human nature? The questions are hard, and the answers are not obvious. But ignoring the problem won't make it go away.
