This essay argues that rational people don't have goals, and that rational AIs shouldn't have goals. Human actions are rational not because we direct them at some final 'goals,' but because we align actions to practices: networks of actions, action-dispositions, action-evaluation criteria, and action-resources that structure, clarify, develop, and promote themselves. If we want AIs that can genuinely support, collaborate with, or even co-evolve with human agency, AI agents' deliberations must share a 'type signature' with the practices-based logic we use to reflect and act.
The argument
The essay argues that these issues matter not just for aligning AI to grand ethical ideals like human flourishing, but also for aligning AI to core safety-properties like transparency, helpfulness, harmlessness, or corrigibility. Concepts like 'harmlessness' or 'corrigibility' are unnatural — brittle, unstable, arbitrary — for agents who interpret them in terms of goals or rules, but natural for agents who interpret them as dynamics in networks of actions, action-dispositions, action-evaluation criteria, and action-resources.
While the issues this essay tackles tend to sprawl, one theme that reappears is the relevance of the formula 'promote X -ingly.' This formula captures something important about both meaningful human life-activity (art is the artistic promotion of art, romance is the romantic promotion of romance) and real human morality (to care about kindness is to promote kindness kindly, to care about honesty is to promote honesty honestly).
Eudaimonic rationality
The essay starts by asking: What follows for AI alignment if we take the concept of eudaimonia — active, rational human flourishing — seriously? The concept of eudaimonia doesn't simply point to a desired state or trajectory of the world that we should set as an AI's optimization target, but rather points to a structure of deliberation different from standard consequentialist rationality. This form of rational activity and valuing, which the essay calls eudaimonic rationality, is a useful or even necessary framework for the agency and values of human-aligned AIs.
These arguments are based both on the dangers of a 'type mismatch' between human flourishing as an optimization target and consequentialist optimization as a form, and on certain material advantages that eudaimonic rationality plausibly possesses in comparison to deontological and consequentialist agency with regard to stability and safety.
The concept of eudaimonia suggests a form of rational activity without a strict distinction between means and ends, or between 'instrumental' and 'terminal' values. In this model of rational activity, a rational action is an element of a valued practice in roughly the same sense that a note is an element of a melody, a time-step is an element of a computation, and a moment in an organism's cellular life is an element of that organism's self-subsistence and self-development.
Key concepts
The essay develops several key concepts:
Eudaimonic practice: A network of actions, action-dispositions, action-evaluation criteria, and action-resources where high-scoring actions reliably (but defeasibly) causally promote future high-scoring actions.
Eudaimonic rationality: A class of reflective equilibration and deliberation processes that assume an underlying eudaimonic practice and seek to optimize aggregate action-scores specifically via high-scoring action.
Support practices: Eudaimonically rational ways to support eudaimonic practices. The work of a good couples' therapist, for instance, is intertwined with but clearly distinct from a couple's relationship-practice.
Implications for AI alignment
The essay argues that many puzzles and 'paradoxes' about AI alignment are driven by the assumption that mature AI agents will be Effective Altruism-style optimizers. A 'type mismatch' between Effective Altruism-style optimization and eudaimonic rationality makes it nearly impossible to translate the interests of humans — agents who practice eudaimonic rationality — into a utility function legible to an Effective Altruism-style optimizer AI. But this does not mean that our values are inherently brittle, unnatural, or wildly contingent: while Effective Altruism-style optimizers may well be a natural type of agent, eudaimonic agents (whether biological or AI) are highly natural as well.
The essay proposes that the cultivation of human flourishing as such — the cultivation of the harmony of a multiplicity of practices, including their resource-hungry support practices — is the cultivation of an adverbial practice that modulates each and every practice. What makes our practices 'play nice' together are our adverbial practices of going about any practice kindly, respectfully, accountably, peacefully, honestly, sensitively.
Virtue-ethical agency
The essay argues that treating core AI-safety desiderata like transparency, corrigibility, and niceness as moral virtues — domain-general, always-on eudaimonic practices — dissolves problems and paradoxes that arise when treating them as goals, as rules, or even as character traits. We want an agent that actively tries to be transparent, and to cultivate its own future transparency and its own future valuing of transparency, but that will not engage in deception and plotting when it expects a high future-transparency payoff.
If this is right, then eudaimonic rationality is not a matter of congratulating ourselves for our richly human ways of reasoning, valuing, and acting but a key to basic sanity. What makes human life beautiful is also what makes human life possible at all.