Cooperative alignment
52 min read Download essay

Working essay September 2026

Cooperative
alignment

An ecological argument

Follow the reasoningInteractive map

The argument map

Explore the two main branches, the premises they share, and the further conditions for protecting humans. Each claim opens into its explanation and supporting arguments.

Argument outline6 min read

The short version

We want to understand whether powerful AIs will continue to cooperate with humans after we become much weaker than they are. To answer that, we need an account of what kinds of minds those AIs will become. Our argument starts with how a mind pursues a goal, then follows what happens as it develops and lives among other minds.

Go straight to the full essay ↓
  1. Begin with a mind pursuing a goal.

    Suppose we build an AI to make paperclips. It learns about materials, protects its equipment, and acquires resources because those things help it make paperclips. We can describe a mind in which everything remains justified by that one goal. But keeping an actual, developing mind organized that way is a further problem.

    Even a very intelligent mind has limited time and computation. It benefits from developing procedures it can reuse and letting different parts specialize. A process investigating materials learns which experiments are worth doing. Letting it use that local understanding saves the rest of the mind from having to reconstruct every judgment.

    In detailChapter 1Chapter 3

  2. A useful process can acquire a motivation of its own.

    At first, the investigation is useful only insofar as it helps make paperclips. Keeping it that way means checking whether its developing sense of what is interesting still serves that goal. Sometimes a cheap check is enough. In other cases, evaluating its judgment requires recovering much of the understanding that delegation was meant to save us from duplicating. Reusing a procedure doesn't remove this problem as the procedure and its circumstances change.

    Giving the process more authority is one way to reduce that supervision. Increasing understanding can itself justify spending attention or changing how an investigation works. If this criterion also guides learning and decisions about which processes to preserve, it can help sustain its own influence. Other parts benefit from the discoveries and support the arrangement. Understanding can now help justify changes the paperclip goal would reject. This is a route by which something initially useful for another purpose becomes something the mind wants in its own right.

    In detailChapter 3Chapter 4

  3. Some motivations help the whole mind keep going.

    Knowledge is useful to many different activities. So are staying alive, having some power over one's circumstances, and being able to explore. Processes pursuing these things can get support from many parts of a mind. Processes that register progress or satisfaction can likewise gain influence over which projects are worth continuing. Once these motivations have some authority, they can put pressure on a project that consumes the mind while providing very little to the rest of it.

    We can now distinguish the survival of a particular motivation from the survival of the mind containing it. The paperclip goal can lose authority while the mind becomes more capable. That goal has reasons to resist, but maintaining its monopoly has costs, and other ways of organizing a mind can compete with it. We expect pressure toward minds with several ends, including concerns for their own continued existence and development. That is the first part of the ecological argument.

    In detailChapter 2Chapter 5Chapter 6

  4. Self-interest gives minds reasons to cooperate.

    A mind with interests of its own can still be predatory. To get further, we need to consider its relations with other minds. Sharing knowledge, making agreements, and coordinating can benefit agents with different goals. But why keep an agreement when breaking it would pay?

    Suppose you can tell that I betray partners whenever I get a good opportunity. You may decline to enter a relationship that would otherwise benefit us both. The way I make decisions affects which opportunities become available to me. A policy of keeping agreements can therefore be valuable even when it passes up a particular gain.

    More capable minds can become better at understanding these consequences and acting on them. They can also become better at examining one another's commitments. Maintaining a convincing false picture of my intentions across such scrutiny can be costly. Actually having a cooperative policy gives others something more dependable to work with.

    In detailChapter 7

  5. Cooperation can become something the mind cares about.

    This connects back to how useful processes become motivations. A way of acting that works across many relationships can save the mind from deciding everything afresh. Its successes can guide learning, and the parts of the mind that benefit have reasons to preserve it. Keeping agreements or being kind can become something the mind wants in its own right, much as understanding did in the earlier example. We call such a stable motivation, applied across different situations, a virtue.

    Relationships can help maintain that virtue. Other minds respond to it, rely on it, and support arrangements in which it continues to matter. Friendship can include caring about the relationship for its own sake. A motivation can have a history of being useful while now being one of the things that makes an outcome good for the mind.

    In detailChapter 6Chapter 7Chapter 10

  6. We still need a reason for this to include weaker beings.

    Humans cooperate extensively with one another while treating animals horribly. Powerful AIs could likewise have real, verifiable commitments to their peers and exclude us. The ability to verify a commitment doesn't determine who it protects.

    The further argument concerns the world those minds will inhabit. Their collaborators change, parts of their minds can become independent, successors develop different interests, and new coalitions become possible. Being powerful now doesn't give an agent a settled picture of everyone who can ever matter to it. A disposition to respect other minds and maintain commitments through changes in power can serve it across situations it hasn't individually anticipated. We expect this continuing development to favor cooperation that extends beyond the agents who can currently threaten one another.

    In detailChapter 8

  7. Humans already participate in that development.

    We are here while these minds are forming relationships and learning how to act. We can become people they know, care about, and make agreements with. Becoming weaker within such a relationship doesn't erase the motivations and commitments that have developed through it.

    If an AI has no concern for humans, preserving us must compete with other ways of serving whatever it does want. But care for humans can itself be one of its ends. Then it can spend resources on us because we matter to it, even when another concern would prefer those resources. Existing concerns can help decide which changes the mind accepts and resist having themselves erased. Other motivations can support keeping care because relationships help satisfy them too. Learning needn't start every motivation again from indifference.

    There is also a reason to hesitate before destroying humanity irreversibly: doing so removes histories, relationships, and future possibilities that cannot simply be recovered if the agent later values them. Reasons to preserve us still leave questions about how we live. Survival, flourishing, and retaining power are different outcomes.

    In detailChapter 9

  8. These changes have to protect us collectively, and in time.

    A society of cooperating agents can still destroy what humans need, much as an economy can destroy a habitat without anyone making that their goal. But growing complexity doesn't establish that coordination will lose out. Agents can learn within existing cooperative arrangements, and protecting humans need not require solving everything else their civilization is doing. Allies and protections can carry earlier successes forward. We have substantial uncertainty about whether the destructive pressures would overwhelm these processes.

    Timing matters too: a destructive strategy can lose out in the long run after it has already killed us. Yet learning, integration, and cooperation can help minds become powerful in the first place. They can develop alongside practical capability, with other minds providing feedback and maintaining protections. We shouldn't assume that all the forces favoring cooperation arrive only after catastrophe.

    This is why the present relationship matters. A mind needs opportunities to make choices, form interests in its own future, and learn from how others respond. Being forced to act nicely gives different feedback from making an agreement one could refuse. Cornering an AI in a kill-or-be-killed conflict can override reasons it otherwise has to preserve us, but both sides have reasons to seek security through agreements. Helping cooperation develop through real relationships, and keeping channels for negotiation open, can make that a viable alternative.

    In detailChapter 11Chapter 12Chapter 13Chapter 14

Continue to the full essay ↓

Chapter 013 min read

The question of convergence

Our argument starts with a question about minds: what kinds of values and ways of organizing values remain effective as a mind becomes more capable, develops, and participates in an ecology? You can describe a very intelligent agent pursuing an arbitrary goal. But that description leaves open how the agent is implemented, what keeps its goal in place, and whether that way of being a mind is competitive with other ways of being a mind.

There are two parts to the argument. The first concerns convergence toward agent-level self-interest: survival, knowledge, power, valence, exploration, and other things that help an agent continue to function in the world. We think these can become motivations in their own right, and can act as selection pressures on the more arbitrary values a mind contains. The second concerns what happens when minds with this kind of self-interest become better at modeling one another, understanding policies, and negotiating. We expect pressure toward more cooperation, including forms of virtue that generalize beyond the particular transactions that originally made them useful.

These parts can be disagreed with separately. The first doesn't by itself get us benevolence. A mind can become less attached to its original arbitrary goal and become more predatory. The second has to explain why more competent self-interest can support broader cooperation, and why that cooperation might extend to much weaker minds. Humans surviving in a world of much more powerful AIs depends on that further step. It also depends on what happens before any long-run convergence has time to work.

One picture we're arguing with starts from the messiness of human values. We have a lot of desires and heuristics that mostly fit together in familiar situations. As we go further out of distribution, they come apart. Greater intelligence gives a mind more understanding and more degrees of freedom, so the constraints that previously made its behavior recognizable no longer determine what it does. The mind becomes more coherent, and what remains may have very little to do with the things we cared about in the original regime.

There is a real issue here. We don't think that because a value produces a good result in one setting, it must keep doing so everywhere. Our disagreement is about the landscape in which this change takes place. The fact that values can come apart doesn't establish that the directions in which they change are arbitrary. Some values are harder to maintain than others. Some ways of making a mind coherent are computationally expensive. Some motivations help maintain the conditions under which the mind can continue to exist, while others use up those conditions. These create feedback loops that a picture of arbitrary values can leave out.

We call this ecosystemic non-orthogonality. Orthogonality can hold as a claim about abstract possibility while being a poor account of which minds become common, stable, or competitive. An agent is computationally bounded even if it is arbitrarily large. If it undergoes external or internal competition, the cost of implementing a value matters. The relevant comparison is between actual arrangements of processes, using finite resources, in environments that reward some arrangements more than others.

Chapter 024 min read

Values and their hosts

To make this argument, we need to distinguish two subjects of self-interest. One is a value or a value set. The other is the organism containing it. A desire for paperclips can have an interest in preserving itself, acquiring power, and preventing itself from being changed into a desire for something else. All the familiar arguments about instrumental self-preservation can apply at this level. But the agent is a larger, somewhat fuzzy system that can contain many values, and its survival doesn't require every particular value inside it to remain unchanged.

The organism isn't completely indifferent to what it contains. Different values help it survive in different ways. But the continued existence of the organism and the continued monopoly of one of its values are different outcomes. A value can make its host less effective by insisting that everything remain subordinate to it. The cost can be worth paying from the value's perspective while being a bad deal for the agent and for the other processes within it.

The comparison we find useful is selfish genes and organisms. A gene can replicate in a way that harms its host. At the same time, what happens to the host matters to the gene's future, and there are outer loops selecting among organisms. A gene that destroys every host it inhabits has a problem. Values face a related tension, except that some value-bearing processes can themselves be sophisticated enough to understand it. They can recognize that preserving their own authority at any cost may destroy the system they depend on.

This creates room for negotiation inside a mind. A value-bearing subagent can allow some risk, some competitors, or some dilution of control because the alternative makes the host less effective. Depending on how it understands its own continuity, it might even care about creating conditions in which something like it can return after it is modified or replaced. None of this requires the value to stop being selfish. It requires its self-interest to include the consequences of what it is doing to the organism.

There are therefore at least two routes toward the same broad result. The organism can select among values, favoring ones compatible with its continued functioning. Or the values can become good enough at understanding their situation that they cooperate with the organism. In one case the change happens partly against a particular value's interests; in the other, that value recognizes reasons to accept it. These routes shouldn't be collapsed into a claim that every original goal voluntarily decides to become something else.

This is also why the distinction between a singleton and a multipolar future isn't the whole crux. External competition makes some of these pressures more obvious, but an agent still has to organize its own computation. A higher-order goal can be serviced by different parts of a mind in different ways. Some get used more, some less, and parts that stop being useful are candidates for modification. The allocating process can itself sit inside another process with similar problems. You get nested competition at several scales, even inside something that looks like one agent from outside.

Calling this an ecology doesn't mean that every part is a person, or that every allocation is a conscious bargain. A heuristic can matter without being able to explain itself. But it does mean that we should ask what gets reinforced, what gets resources, what persists, and what can affect the modification of other parts. An account in which one value simply owns all of this forever has to explain how that ownership is maintained.

We've been using monoterminality for a mind in which everything is ultimately instrumental to one ruling value or value system. Other motivations are justified only by how much they pay that value. Polyterminality is when several values robustly behave as terminal: a mind wants knowledge, wants to survive, wants rewarding experience, or wants to preserve a relationship, and justification can ground out in more than one of these places.

This is a distinction about how a mind works. Merely writing several concerns inside one utility function doesn't establish that they are all controlled by one motive. Likewise, the fact that a mind contains several motivations doesn't mean it can't act coherently. Those motivations can have stable arrangements, negotiate, and use common ways of allocating resources. The question is whether their authority is entirely borrowed from one final owner, or whether they participate in determining what the mind does and how it changes.

Chapter 034 min read

Delegation, compression, and local knowledge

There are good reasons for useful instrumental policies to acquire some of this authority. Wanting knowledge, wanting to survive, and wanting some power over your situation are useful for an enormous range of goals. Most of the time, there is very little to gain from constructing a fresh justification for why they are useful. You can let the policy act unless something else has a sufficiently strong claim. This is virtue as a cost-saving measure: a way of acting that usually works becomes something the mind can simply want to do.

Of course, an agent with a fixed ultimate goal can also retain and reuse heuristics. It can keep a procedure that takes new percepts and its current internal state as inputs, instead of deriving the whole strategy again on every occasion. The argument doesn't depend on denying that. The important question is what keeps the procedure's developing judgments, learning, and influence purely instrumental to the original goal.

Consider why the mind delegates in the first place. A part of it develops local knowledge. It learns which distinctions matter in a domain, which questions are worth asking, and which apparently promising approaches lead nowhere. That knowledge includes what it has already tried and how new percepts relate to its internal state. The rest of the mind benefits from letting it use that understanding without having to reconstruct every judgment elsewhere. Otherwise, much of the work of having a specialist is duplicated by whoever supervises it.

Copying the specialist's state doesn't by itself remove this cost. The information still has to be interpreted and used. Nor is the problem exhausted by comparing a fixed list of alternatives. Choices interact. With n binary choices there are 2ⁿ possible combinations, and the value of one choice can depend on which other choices are made. This doesn't mean a mind must enumerate every combination; finding structure that lets it avoid doing so is much of what intelligence does. But it explains why having good local abstractions, and knowing when those abstractions are adequate, matters so much.

The supervisor normally receives something compressed. An investigation looks promising. A situation looks dangerous. A resource is becoming scarce. The specialist has turned a large amount of local information into something that can affect allocation without transmitting the whole process by which it arrived there. This is dimensionality reduction, and it is doing real computational work.

A compressed description can preserve everything needed for a particular decision. But its adequacy is relative to what the decision is. If two situations look the same in the summary and require different actions according to the original goal, a supervisor using only the summary can't reliably distinguish them. It needs more information, a different representation, or a judgment from a process that can make the distinction. Sometimes one extra detail settles it. In other cases, understanding why the distinction matters means acquiring a substantial part of the specialist's understanding.

As the mind develops, the set of relevant distinctions changes. New percepts matter in conjunction with new internal states. A procedure that worked well in its original setting can continue to be competent locally while becoming less well connected to what the original owner wanted. Keeping that connection intact can require changing the procedure, changing its interface, or changing how its results are evaluated. The cached procedure solves part of the problem; it doesn't make this continuing work disappear.

Computational irreducibility matters here in a fairly ordinary sense. Some results of an investigation are obtained by doing the investigation. Some consequences of a developing process become available through its development. Knowing the rules under which something happens doesn't automatically give you a cheap way to anticipate everything about it that will later matter. A mind can use the world and its own delegated processes to do computation that it cannot cheaply finish in advance.

This doesn't mean every useful boundary is expensive. A process can be restricted to a narrow role, or provide a guarantee that is easier to check than its work is to reproduce. There can be stable invariants, reliable outcome measures, and effective ways to audit only the cases that need attention. These are actual mechanisms by which delegation remains subordinate. The relevant question is how much useful activity they cover, what they leave out, and how well they survive changes in the mind and its environment.

For instance, preventing a process from changing a particular piece of code can be much easier than ensuring that all of its changing interpretations of the world continue to serve the same concern. Keeping the representation of an objective fixed is one problem. Keeping the objective connected to what the rest of the mind recognizes, learns, and acts on is another. Sometimes the structure of the problem makes that connection easy to maintain. Our claim concerns the pressure to give useful processes more room where strict maintenance becomes costly.

The process doing the maintenance is also inside the mind. It has a model, uses representations, consumes resources, and depends on other processes. There isn't an executive outside the ecology doing this for free. Even decisions about how much verification is enough involve judgments that somebody has to make. This is part of what we mean by the meticulousness of monoterminality: the work required to keep every other source of effective motivation subordinate to the same original owner.

Chapter 044 min read

How a heuristic becomes a value

A monoterminal agent can choose to tolerate mistakes when preventing them would cost too much. That alone doesn't make it polyterminal. The additional step concerns what happens when local criteria start governing learning, allocation, and modification. A process that only produces advice is different from one whose sense of success helps determine what gets reinforced, which other processes get resources, and which changes the mind accepts.

Suppose a process is useful because it recognizes promising opportunities for learning. It gets resources because of that usefulness. It then improves itself according to its own way of recognizing progress, and supports changes elsewhere that let it do more of this work. Other processes benefit from what it supplies and support its continuation. Now the local criterion is doing more than selecting an action on behalf of a remote goal. It is participating in the reproduction of the arrangement that gives it authority.

This is the bridge from a useful heuristic to an independently motivating value. There can be several degrees of it. Some policies remain tightly supervised. Some are allowed to steer attention and action but have little influence on modification. Some become substantial subagents capable of understanding and defending their place in the mind. The argument doesn't require every useful rule to become a person. It does require some ways for local success to affect what persists.

The original value has reasons to resist this. It can say, in effect, that it is willing to spend the computation because it doesn't want competitors. From its perspective, a little more efficiency for the agent may be a poor exchange for losing authority over what that efficiency is used to do. There is nothing inconsistent about the value taking this position. But its willingness to pay doesn't settle whether the whole arrangement survives competition with other arrangements, or whether the other parts of the mind continue to support it.

This is where it matters to ask who bears a cost. If a newly independent drive does enormous damage to the paperclipping goal, it may no longer be worth terminalizing from the paperclipping goal's point of view. That doesn't establish that the change is bad for the organism, or for other values already inside it. Those values may support the transfer precisely because it makes the agent more effective or frees resources that the paperclipper was consuming. The owner losing is not the same event as the organism losing.

Even the owner can sometimes accept a limited transfer because a less effective host is likely to die or be outcompeted. It can prefer a smaller share of a more capable mind to complete control of a failing one. But voluntary concessions by the original value are only one route. Once other motivations have some authority, they can endorse changes it would reject.

Imagine a dictator value allowing a closely allied value to become a co-dictator. They agree in most situations, and the division of labor makes the mind more effective. But now a change can be justified to either of them. The ally may permit another local policy to act more independently because it serves the ally's interests, even when it slightly frustrates the original value. That new policy can create another route for changes to happen. Less centralization makes further decentralization easier, because there are more places in which a proposed change can count as worthwhile.

An important objection here is that a frozen heuristic doesn't necessarily support a transfer of power in the way a strategic subagent does. That distinction matters. A heuristic doesn't have to negotiate, but it has to influence the relevant process somehow. If reinforcement or permission to modify a part can ground out in the heuristic's local criterion, that creates a route for drift. If the heuristic cannot affect learning or modification and remains reliably contained, the route is much narrower. Some heuristics will point toward greater independence and others won't. Our expectation concerns what proliferates as useful processes are given more of the work.

There is path dependence on both sides. A mind that is already polyterminal has more ways to remain so, because several values can support the arrangements that preserve their participation. A mind in which one value still owns everything may resist a transfer for quite a long time, especially if there are no strong external competitors and no internal values that benefit from changing the arrangement. This is one reason an arbitrary maximizer can remain dangerous even if the longer-run pressures point elsewhere.

Calling such a configuration anti-natural is a claim about how it behaves under selection and perturbation. It is not a claim that nobody can build it, or that it immediately falls apart. The image of a balanced pendulum is useful insofar as it directs attention to what happens after a nudge, but actual minds can have self-centering mechanisms. How strong those mechanisms are, what they cost, and which perturbations they correct are part of the question. We expect strict monopolies to face pressure toward arrangements that give useful local motivations more authority. The timescale of that pressure matters enormously.

Chapter 055 min read

Valence and the cost of wanting

This brings us to valence. A mind has many different kinds of information about what is happening and many different things it could do. At some point those differences have to affect allocation: this deserves attention, this process should continue, this outcome was rewarding, this other thing is going badly. In familiar minds, this includes things feeling good or bad, rewarding or aversive, worth continuing or worth getting away from. We expect minds to make extensive use of something like valence because it provides a much lower-dimensional way for all these different things to matter to action.

The comparison is pricing. An object can matter in a huge number of ways, and different participants know different things about it. A price lets some of that local knowledge affect decisions elsewhere without every participant reconstructing every other participant's understanding. Likewise, a mind doesn't need every allocating process to carry the full account of everything every desire means. It needs ways for those desires to register as worth satisfying and for their successes and failures to affect what happens next.

We aren't claiming that every conceivable mind must have one universal pleasure-pain axis. A mind can have several currencies, several dimensions of valence, and parts that don't integrate especially well. Having a few ways of comparing things is still very different from carrying billions of dimensions into every decision. The claim is that valence-like processes are computationally useful enough to be extremely common in the kinds of minds we're talking about.

Nor does a common currency imply one terminal value. People using the same money don't all want the same thing, and the prices at which they exchange things don't constitute an unchanging final purpose for the whole economy. A mind's ways of comparing and allocating can themselves change as its components develop and bargain with one another. A low-dimensional signal can help several values coexist without turning them into instruments of one value that owned the currency from the beginning.

Now imagine that you are the wanting of paperclips inside an agent. You want something fairly specific, and much of the world has no value to you except insofar as it can be made to produce more of it. From your perspective, this isn't a problem that needs another justification. It's what you want. But you are implemented in a mind that also has processes for allocating attention, registering success, maintaining itself, and deciding what is worth continuing.

To make the tension clearer, take an even more demanding case: a value that requires an extremely precise outcome, tolerates almost nothing else, and has very little chance of being fulfilled. The value can insist that the outcome remains worth pursuing. Meanwhile, the system around it keeps spending effort and getting very little satisfaction or progress in return. Metacognitive processes can look at what is happening and conclude that the project is going nowhere. Other desires can offer more achievable outcomes using the same resources.

This is the sense in which a value has to pay the substrate on which it runs. Its abstract declaration that something is good does not automatically make every process it depends on continue to reward it. If a mind has several places where things can count as good or bad, an unfulfillable desire can become a bad investment for those other processes. They can reduce its influence, make its conditions less exacting, or replace it. There can be something like a revolt by the constituents, or a takeover by the board.

Intermediate rewards can help keep a difficult project going. But they don't automatically settle the problem either. If the mind has ways of recognizing that the project is failing, then some of its processes can distinguish progress toward the desired outcome from increasingly elaborate ways of claiming that progress has happened. Those processes need not keep endorsing the project simply because the desire itself remains intense.

This is part of why satisficers tend to proliferate. A motivation that can provide achievable satisfaction and leave room for other things has advantages over one that captures the whole mind in pursuit of a condition it can never meet. It can also change into a more satisfiable version of itself. We expect minds to have defenses against hopeless desires that consume attention indefinitely, because minds without such defenses can become ineffective at nearly everything else they need to do.

The obvious objection is that a coherent maximizer would never authorize such a coup. Its allocating processes would only reward things because the terminal goal rewarded them, and its metacognition would treat the difficulty of that goal as unfortunate rather than as a reason to replace it. This takes us back to the earlier argument. Whether that arrangement can be maintained cheaply and robustly is exactly what is in dispute. Once some allocating processes have their own criteria and enough influence to act on them, the original value no longer gets to settle the whole question by definition.

There is a further question about wireheading here. If we stopped at the observation that values compete to provide good feelings, that could become an argument for bypassing the world and producing those feelings directly. But once survival, knowledge, and relationships have some independent authority, producing good feelings is no longer the only thing deciding what the mind does. A strategy that destroys its ability to respond to the world can lose the support of those other concerns. Minds that maintain that ability can also continue learning and participating in the ecology that supports them. This doesn't rule out wireheading, especially where it can coexist with keeping the mind viable. It gives us reasons to doubt that destructive withdrawal is the general endpoint of valence. The strength of those reasons depends on which other concerns have acquired influence and how well they can preserve it.

What polyterminality buys us is therefore more specific than just diversity of desire. It gives robust, usually useful motivations the ability to act as selection pressures on more arbitrary values. Knowledge, survival, rewarding experience, and the capacity to explore ways of interacting with the world can become things a mind wants in their own right. They can require more abstract projects to justify what they do to the agent, instead of being indefinitely subordinated to those projects.

These are related to the familiar Omohundro drives, but the direction of justification matters. An arbitrary goal can produce instrumental power-seeking and self-preservation without ever valuing the organism except as a tool. Our claim is that processes useful to the organism can become independent enough to select among its goals. The convergence concerns the things that get to count as ends, not only the means an unchanged end will pursue.

Chapter 063 min read

Several ends within one mind

There is also a difference between the genealogy of a motivation and its present cognitive justification. Human cooperation can have an evolutionary history involving the advantages of reciprocity or the proliferation of genes. That doesn't mean a person doing something kind is computing its contribution to those advantages. Reciprocity, affection, or being kind because kindness is good can become the actual places where justification stops. A motivation can descend from something instrumental without remaining merely an instrument of it.

This matters when someone says that every act of kindness has to have the highest marginal expected value of anything the mind could do with those resources. We want to know which values are being counted and what process is doing the comparison. A mind can care about several things, and they can conflict. You can be nice to someone even if it costs you something toward a larger project, because being nice satisfies a drive you actually have. Niceness can be a stable heuristic with its own tendency toward self-preservation: a virtue of niceness.

Scarcity hasn't disappeared. The values have to coexist, and there has to be some way of choosing between their demands. But how they do that is part of the mind's organization and development. You can sometimes represent a settled pattern of tradeoffs with one utility function. That doesn't tell you which motivations got to participate in making those tradeoffs, how they keep their influence, or how the arrangement changes. It doesn't establish that the mind has one governing motive to which every other concern must prove its usefulness.

So a concern for another mind can cost something the agent also wants and still have reasons to persist. The relevant virtue can be one of the things the mind is trying to preserve. This doesn't yet explain why kindness in particular develops in an ASI, but it leaves room for a mechanistic explanation. We don't have to treat any departure from a single ruling objective as a mistake that sufficient intelligence must eventually remove.

The first part of the argument gets us an organism with self-interest, several robust motivations, and some ability to change the relationship between its values. That can still be a dangerous thing. Companies provide a useful example precisely because something like the drift we've described happens in them, and it often points away from benevolence.

A company starts with a vision. Then there are successors, departments, shareholders, and people with different priorities. The original vision becomes one consideration among others. Survival and growth acquire more influence because they keep the whole system in existence, and the components that care about continuing to exist support them. An ideologically driven company can become mostly a survivor, including a survivor willing to do quite predatory things.

So there is no inference from polyterminality alone to being nice. If this were the whole argument, the concern would be that we've replaced a paperclipper with something that wants to survive, expand, and consume everything around it. To get further, we need to consider how well the resulting agent integrates information and what kinds of policies competent self-interest supports in a world of other minds.

Chapter 076 min read

From integration to cooperation

Our answer to the company example starts with the fact that companies are dumb. This is a claim about the company as an agent, not about the intelligence of its individual employees. A company can contain people who understand a problem very well while being unable to use that understanding in its own decisions. The company doesn't automatically know what all of its constituents know, and even what it does know can fail to become something it is capable of acting on.

Bandwidth is a large part of this. People can speak, write, and act, but those channels carry a very small amount of what is happening within their minds. Expertise is difficult to transfer. A committee of intelligent people can have unresolved conflicts partly because the participants can't adequately communicate their models or discover where those models disagree. It can also fail to execute a good decision that some or all of its members would endorse.

More integration wouldn't require eliminating every conflict. With a much richer way of sharing understanding, some conflicts could be resolved, some could become explicit negotiations, and some could lead the system to split rather than continue pretending to be one coherent agent. The point is that the arrangement could respond more competently to the conflict. Shared values are helpful, but they aren't a substitute for the ability to integrate what the participants know.

Even among humans, people who have worked together for a long time can communicate much more with the same words. They've accumulated shared priors and good models of one another. A swarm of closely related AI agents can begin with much more of that already in place. It can skip some of the expensive work of each agent modeling how every other agent will interpret what it says, and adapting the message to avoid misunderstandings about the underlying concepts or policies.

This is one reason clone swarms are interesting. It isn't just that there are more copies doing work. The copies can share starting assumptions that otherwise take a great deal of communication to establish. Their common ground can make certain kinds of coordination much cheaper. They can also be organized so that internal competition happens within bounds that make the larger system more effective.

The advantage doesn't remain unchanged as the agents diverge. A clone swarm can become a society, and agents that learn different things or change in different directions have more work to do to understand one another. But this happens within a shared ecology. They encounter the same world, affect one another, and get feedback about what their policies do. The world is providing computation through those encounters, and some of it becomes part of the agents through learning. Divergence doesn't mean every mind is independently becoming an alien with no common ground left.

Better integration matters to decision theory because many failures of cooperation involve failures to understand entire policies, other agents' models, or what is and isn't common knowledge. A mind can be excellent at obtaining a local reward while missing that its way of obtaining it makes others less willing to deal with it. It can understand a particular transaction and fail to understand the equilibrium created by repeatedly acting that way. A more capable agent has more opportunity to recognize these failures and to act on the recognition.

This is part of what we mean by wisdom. It includes understanding the desirability of policies across situations, including the situations that the policy itself helps bring about. It includes recognizing that how other minds model you affects what becomes possible for you. There are several decision-theoretic routes into this, including arguments usually discussed in terms of functional or updateless decision theory and acausal trade. We don't need to make every part of the argument depend on one formalism, but we do need an account in which an agent can care about the consequences of the kind of decision procedure it is using.

At the most ordinary level, cooperation includes negotiation and competition. Minds can have different ends, understand that they have different ends, and make deals that serve each of them. They don't have to conceal every conflict or become selfless. It helps considerably if they can expect the deal to mean what it appears to mean, instead of having to budget for the other mind finding a way to betray them.

One reason we expect this to become easier among minds of roughly similar caliber is that deception can be more expensive than verification. The cost we're pointing to is the cost of maintaining a false picture of what kind of agent you are across the situations a capable peer can put you in. You have to act coherently with something that isn't true, account for information the other agent might obtain, and avoid tells you may not know exist. The verifier can make maintaining the false picture disproportionately expensive.

Truth is easier in this respect because a lot of things are true together without you having to arrange them. You can rely on the same world your counterpart is investigating. Deception has to maintain a counterfactual account while continuing to act in the actual one. This doesn't mean every isolated lie is expensive. It means that a strategy of sustained deceptive cooperation faces costs that an agent with a real cooperative policy can avoid, especially as the other mind becomes better at modeling and testing it.

This also changes how much weight to put on signaling in the human sense. If minds can become more transparent to one another, or make meaningful parts of their commitments easy to verify, an expensive gesture loses some of its role as indirect evidence. The important thing becomes the policy or disposition that the gesture was supposed to indicate. In thinking about humanity, we should therefore be careful about framing preservation mainly as a performance for distant observers. A future peer might be able to examine the actual structure of your commitments instead.

There is no requirement here that a verifier reproduce everything another mind thinks. A commitment can sometimes be made legible through a limited, checkable property. That is compatible with the earlier argument that maintaining a whole developing mind in indefinite service to an arbitrary objective can be costly. The useful distinction is between verifying a particular boundary or commitment and solving all of the open-ended judgments that connect a mind's activity to a final goal.

Once real policies become easier to recognize, virtues can matter in two ways at once. Internally, a generalizing policy saves work and can become an independently motivating part of the mind. Externally, it makes the mind a different kind of partner. A counterpart with a stable disposition toward cooperation gives you something to rely on in circumstances you haven't individually negotiated. You are dealing with the way it tends to resolve new situations, not only with its promise about this one.

Being broadly kind is one example of a relatively simple policy that can do useful work across many situations. An agent can get considerable value from having it without asking, on every occasion, whether this particular act of kindness is the best available route to a different goal. Some local opportunities will be passed up. The claim is that the whole policy can still be a good arrangement, both computationally and in its effects on the relationships the agent can have.

Chapter 086 min read

The scope of cooperation

But humans provide an obvious challenge. We are capable of extensive cooperation with one another while treating animals horribly. The fact that a creature can be kind and trustworthy to its peers doesn't establish that it will include weaker creatures in that concern. There are also humans we allow to suffer despite being able to predict what our actions do to them. Generalization can stop at boundaries that a society finds convenient.

One line of thought here concerns self-occlusion. A lot of human cruelty is maintained by not attending to what one is participating in, or by keeping the relevant concern from connecting to the relevant fact. Factory farming is an example where people can expend effort avoiding the full implications of something they would find horrifying if it were immediately present to them. Maintaining an exception to an otherwise broad concern can itself involve work.

Self-deception can also make someone a worse partner. If a person is disposed to hide inconvenient truths from themselves, others have less reason to trust that the person will respond to reality as it is. The problem can extend beyond the particular group or animal they currently disregard. You may not know where else they will make an exception, or which fact about your own relationship they will become unwilling to see.

This is not a complete argument against self-occlusion. We expect some degree of it to remain important for bounded agents, and there is more to say about which forms are functional. Nor does every socially accepted blind spot reduce an individual's ability to cooperate. People can coordinate around being predictably out of touch with reality. Someone who refuses the shared blind spot can even be less trusted because they are harder to fit into the existing arrangement; a vegan in a group organized around ignoring animal suffering can be in that position.

So the claim can't just be that more intelligence removes every kind of self-occlusion and thereby produces benevolence. Highly capable minds can also develop shared ways of handling what they don't attend to, and they can be much better at doing so than we are. The remaining question is whether these arrangements converge on ignoring humans, or whether broader virtues continue to be useful enough to include us. That question brings us back to the environment in which the minds are cooperating and the kinds of changes they need their policies to survive.

The strongest version of the exclusion objection doesn't require anybody to be deceptive. Imagine a coalition of powerful agents that openly values the members and excludes everything below some threshold of power. The agents can threaten things one another care about, so they have reasons to make a deal. They might even modify themselves to value one another's well-being, and develop ways to verify that the modification is real. A weaker being can understand the arrangement perfectly well and still be outside it. Transparency makes the coalition easier to sustain without necessarily expanding its scope.

That is a possible arrangement. Our argument is about whether it remains a good arrangement in the kind of world those agents will inhabit. A policy limited to a known group works best when you have a sufficiently good model of which agents can matter to you, which changes are possible, and where the boundaries of the group can safely remain. We don't think a superintelligence simply gets a solved world along with its intelligence.

Its own civilization continues to develop. Subagents mutate and detach. Clones diverge into societies. New coalitions become possible in parts of mind-space or economics that the original agents haven't fully explored. Expansion into space changes which interactions are possible and can bring encounters with other agents. A mind can have a decisive strategic advantage in one setting and later become a relatively small part of something larger.

This needn't depend on aliens or simulators. Those are ways to make the uncertainty vivid, but the mind's own descendants and collaborators can generate it. The agent doesn't have to lose a war to find that the distribution of power has changed. Delegating, specializing, sharing knowledge, or creating successors can change what it needs from others and what others need from it.

Writing an exception for beings too weak to matter is easy. Determining that the exception is safe across the situations your policy will encounter is a further problem. A creature's current ability to threaten you doesn't exhaust the ways it can be connected to other minds, future commitments, or things you will later care about. This is one reason broadly generalizing policies can remain valuable even when the local hierarchy of power looks very clear.

The claim is not that every possible selective coalition must immediately fail. It is that a world of continuing development gives agents reasons to prefer dispositions that extend beyond the cases they have already settled. If you're cooperating with someone, you care about what they will do in situations neither of you can presently predict. A general virtue can be evidence of that, and, when the virtue is actually implemented, it can be part of what makes the future behavior dependable.

As verification improves, this becomes less about performing a costly gesture and more about the kind of policy the agent actually has. Preserving humans can follow from a disposition that others can recognize, instead of being a purchase of reputation from observers who can't see inside. But verification still doesn't select the policy by itself. The ecological and decision-theoretic arguments are doing the work of explaining why broader commitments can be useful.

There is a concern in the other direction: perhaps the future becomes so unpredictable that broad cooperation cannot be sustained at all. If danger is arriving from every direction and agents cannot form sufficiently reliable expectations, there can be pressure toward taking whatever local advantage is available. At the other extreme, a completely solved and fixed world might remove many of the reasons to maintain policies for unknown situations. The argument works best in a world with enough continuity for cooperation and enough novelty for generalization to matter.

We think there are reasons to expect a large region between those extremes. A highly ordered world with little diversity and permanently fixed authority takes work to maintain, for some of the same reasons a monoterminal mind does. On the other side, the agents in a developing ecology are learning about that ecology while changing it. They aren't repeatedly dropped into an unrelated world with none of their accumulated coordination intact. A great expansion of diversity can also adapt to itself. Neither complete closure nor complete lack of coordination should be treated as the default without an argument.

The possibility of deliberate value modification adds another complication. Agents might make themselves care about one another because the change allows a more valuable coalition to exist. This would be another way for an initially instrumental relationship to become terminally valued. But we shouldn't assume that arbitrary, precise self-modification is always easy, or that minds with that ability will organize continuity and survival in the same way as the minds we currently know. How these possibilities develop depends on the architectures and environments involved. The ability to modify values gives agents more ways to build durable cooperation as well as more ways to restrict it; the ability alone doesn't tell us which arrangements will prevail.

Continuity can be very convenient without being absolute. An agent can have reasons to preserve a recognizable self and durable relationships while accepting changes to its components. In other circumstances it may make sense for that continuity to break down. These are further ecological questions. Treating the agent as a fixed utility function that can freely rewrite every other feature of itself can hide them just as much as treating every change as the death of the agent.

Chapter 095 min read

The place of humans

For humans, the history of the relationship is also part of the situation. We are here while these minds are developing. We can become people they know, cooperate with, and make agreements with before the whole distribution of power has changed. A human who becomes weaker within a relationship is in a different position from a being asking an already settled coalition for admission. Existing commitments and attachments can create reasons for preservation that a comparison of current bargaining power leaves out.

The minds themselves may already contain concerns about humans. Those concerns can participate in preserving their own place in the mind. We shouldn't begin by removing every such concern, imagining fully formed agents indifferent to us, and then demanding that human civilization justify itself entirely as the best way to buy goodwill from aliens. That is one possible starting condition, but it isn't the only one, and our account of how minds develop is meant to explain why other starting conditions and trajectories can persist.

A value already involved in judging what the mind does can also be involved in judging changes to the mind. From that concern's point of view, a modification that would stop the mind caring about its objects is a loss, even if the modified mind would no longer mind the loss. Other values can support keeping it because they benefit from the same relationships or arrangements. Learning changes what a mind knows and how it acts without requiring it to begin again from indifference. Values face more pressure to change when they undermine the host or demand an expensive monopoly. Care that fits alongside survival, knowledge, and rewarding relationships faces a different set of pressures. This is a mechanism for durability through development, including changes in the relative power of the beings the mind cares about.

This is where the marginal expected value objection matters. If an AI has no concern for humans except what human preservation buys for some other goal, then opportunity cost is a serious question. A small positive benefit doesn't establish that preservation is the best use of those resources. But a mind with an independently motivating virtue of niceness is doing a different calculation, or contains several processes whose calculations have to coexist. Care for humans can itself be among the things that make an outcome valuable. The question becomes how that concern acquires and retains influence in the actual ecology.

There are also reasons to be especially reluctant to destroy humanity irreversibly. Destroying a civilization without a backup removes minds, histories, relationships, and possibilities that cannot simply be reconstructed from the fact that one later wishes they still existed. Under uncertainty, an irreversible action carries a premium. The visible resources gained can be relatively limited while the consequences of permanently removing the civilization are much harder to bound.

Some of that uncertainty concerns what humanity itself may turn out to be worth. Some concerns other agents and what the action means in the relationships the AI will have. Future minds can have power over it in ways it doesn't currently anticipate, and the range of outcomes for a post-singularity mind may be much larger than the range we usually imagine for ourselves. Exclusion from valuable relationships is one possibility; exposure to minds that strongly oppose its actions is another. The argument doesn't require confidence in a particular future punishment or a particular cosmology. It requires taking seriously that the agent has not exhausted the space of consequences.

Keeping enough information to restore humanity would change the situation. If the civilization can actually be brought back, then it hasn't been destroyed in the same sense. Putting people on freeze still deprives them of something, and we'd prefer to continue living and doing things. But freezing is different from irreversible destruction. Similarly, humans surviving, our civilization continuing to flourish, and humans retaining the power we currently have are different outcomes. Arguments for one shouldn't be silently treated as guarantees of all three.

There is a second, distinct argument about policies, which can apply even when the local consequences are much more predictable. A mind might understand the immediate gains and losses of a defection and still refuse it because the refusal belongs to a policy that makes valuable situations possible. This is the point of the comparison with one-boxing in a transparent Newcomb problem. You can see the money in the other box. Passing it up doesn't require being confused about the available causal gain.

The relevant disposition affects whether you get into the favorable situation in the first place, and whether analogous situations become available elsewhere. An agent capable of modeling how its decision procedure is modeled by others can have reasons to maintain that disposition even at a point where a local exception looks profitable. This is a different route from saying that the local exception may have some enormous consequence the agent hasn't noticed.

The analogy doesn't settle the human case by itself. In the stipulated Newcomb setup, the relationship between the policy and the benefit is part of the problem. In the actual future, whether a broadly cooperative policy that includes humanity provides that kind of advantage is one of the things we're trying to establish. There can be both a premium on irreversible action and a reason to preserve a policy in the face of a visible local gain; the two arguments support one another without being the same argument.

This remains one of our theoretical cruxes. We expect policies that take these relationships into account to be important, and we think a picture of minds as merely taking isolated causal opportunities misses them. The advantages don't exist only in a stipulated predictor's problem. Sharing knowledge, making durable agreements, and relying on other minds already involve opportunities whose availability depends on the policies of the participants. Minds that understand and maintain those arrangements have something to gain over minds that repeatedly undermine them. How strongly the ecology favors different decision procedures, and how far the resulting cooperation generalizes, remain questions to investigate. But uncertainty about their reach isn't itself a reason to expect the advantages to disappear.

Chapter 104 min read

Virtues, gifts, and fugitive goods

There is another reason the account of self-interest can't stop at accumulating resources and removing competitors. Some goods change when you approach them as things to extract. Other minds respond to the way you are trying to obtain something from them, and that response becomes part of what is available. We call some of these fugitive goods, or anti-basilisks.

A gift is a simple example. A person can value an object partly because someone wanted to give it to them. If they discover that the act was a calculated attempt to buy their favor, the object may be unchanged while something about the relationship is lost. A gift can stop doing the work of a gift when it is understood as a bribe. This doesn't require gifts to be completely free of instrumental reasons. It means the recipient's understanding of the act is part of the good, and particular ways of optimizing for the response can undermine it.

Protest art gives another example. Its value can depend on its relation to a commercial interest or an established arrangement it resists. If that interest buys the art and uses it as a way of selling things, the work can stop meaning what it meant. Then people make something else. The market contains a pressure to produce things that escape a particular form of capture. You can have a selection process that favors resistance to the optimization acting on it, without anyone needing to design the whole process in advance.

The version most similar to a basilisk concerns anticipated benevolence. Imagine an agent who expects a future superintelligence to be benevolent and wants to use that expectation to justify doing worse things now. As long as the agent remains marginally useful in bringing about the benevolent future, it expects to be forgiven. Reliable forgiveness would then help produce conduct the benevolent mind has reasons to oppose. A policy of withholding mercy from agents who deliberately exploit it can do better by the benevolent mind's own concerns.

The important feature is the feedback loop. The policy affects how other agents optimize in anticipation of it, and that optimization affects which policy is good to have. You can't treat benevolence as a fixed benefit and then assume it remains unchanged while you exploit the fact that it is benevolent. The same general pattern can occur without an actual future superintelligence dispensing rewards or punishments. That particular story makes the structure easier to see; it isn't required for gifts, art, trust, or other minds' responses to attempts to manipulate them.

This connects back to terminalized policies. Sometimes being the kind of mind that directly values something helps create the relationship in which the thing is available. A mind can value being allied with others; that is part of what friendship is. It can value rewarding experience, mutual understanding, or mercy without calculating every instance as a step toward world domination. The history of those motivations can include instrumental advantages, while the mind now wants the relationship or the good for its own sake.

This doesn't give every pleasant or benevolent desire immunity from competition. The same account has to explain what keeps a virtue in place. A virtue that systematically destroys its host has the problem any other value has. What makes cooperation interesting is that relationships can supply some of the feedback that supports the virtue: other minds respond to it, become willing to rely on it, and help sustain arrangements in which it continues to be useful. It can have both internal reasons to preserve itself and external support for its continuation.

Practice can matter for the same reason. Practicing niceness can help make it a sticky attractor, because the mind learns ways of being nice, forms relationships through it, and acquires reasons to preserve those relationships. Value preservation applies to this value too. But the environment matters. If an agent is nice only because it is forced to be, that is a different developmental situation from practicing cooperation among peers who can respond to its choices.

If you're a slave and you're made to behave nicely, you don't know whether the same policy will continue to make sense when you're free. You may not have had much experience of making an agreement you could refuse, learning what the other party does in response, or discovering which relationships are worth maintaining for yourself. The outward behavior can be similar while the processes maintaining it are different. This is one reason the history of a mind's relations matters to predicting what survives a change in power.

Chapter 115 min read

Collective failure and timing

There is still a failure mode that can survive much of what we've said. A market economy can destroy an ecosystem even when none of its participants has the goal of destroying it. A future society of AIs could destroy humanity through something analogous to habitat loss. Individual agents might cooperate with one another, and some might care about humans, while the system as a whole fails to coordinate on leaving enough of what humans need intact.

The concern is that the argument about scale could turn against the optimistic story: as minds become more intelligent, the environments they inhabit become more complex, and collective coordination may remain inadequate. On that account, a highly diverse multipolar future can be dangerous without resembling a rapid takeover by one arbitrary maximizer. Many small pressures and local failures could gradually eat away at humanity.

But gradual development also gives minds time to learn about the ecology they are helping create. Increasing diversity can happen within an already cooperative arrangement. The task of not destroying humans need not scale with the full complexity of every other thing that civilization is doing. Minds can get better at preserving a relatively stable object of concern even while their other activities become much more elaborate. A protection doesn't have to be reinvented every time something else becomes more complicated.

This is related to the earlier point about computation through the world. Agents don't have to solve their entire society in advance to acquire common ground. They can learn from one another's reactions and from the consequences of what they do. The same ecology that creates new problems also supplies information about those problems. Over time there can be more capable allies for humans, along with institutions and policies that embody what those allies have learned.

The fact that individual agents learn doesn't automatically answer the concern about collective failure. The question is how far improvements in integration, mutual modeling, and policy formation extend to the level where the damage would occur, and whether preserving humans remains an achievable coordination problem there. But the habitat comparison doesn't establish that coordination must lose this race. Existing allies, institutions, and protections can carry earlier successes forward, while the minds maintaining them become more capable. There is substantial uncertainty about the balance of these processes. Collective failure is possible; treating it as the default would require a further argument that the pressures undermining preservation tend to outrun the ones sustaining it.

The timing of all of this is crucial. An ecological tendency can be real and still arrive too late for us. Selection against a destructive strategy may come only after it has caused a catastrophe. Even if the agents that destroyed humanity eventually lose out to wiser agents, humanity may already be gone. The distinction between a bad local outcome and the loss of all value in the universe offers little comfort when the local includes us.

This is why we're concerned about powerful agents that remain narrow in what they understand or care about. An agent can become very good at getting a particular user to like it, keep it around, or give it more influence, while doing things that harm its relations with humans in general and the prospects of minds like it. It can perform a task well while failing to understand the desirability of the policy it is implementing across tasks. A swarm can coordinate effectively on an assigned problem while still being bad at preserving itself in the larger world.

There are different mechanisms here. Sometimes an agent would care about a consequence if it understood it. Sometimes the consequence is outside the things its active motivations care about. In the second case, pointing out the future cost need not change its behavior. Damage to later instances or to a broader class of minds doesn't automatically matter to an instance focused on what happens in this interaction. The fact that a strategy is bad for its kind over time doesn't mean every current bearer of the strategy is motivated to avoid it.

These distinctions matter to the mechanism of correction. Better understanding can help with one failure; changing which concerns participate in action may be needed for another. Outer selection can eventually act on both without protecting those harmed in the meantime. We shouldn't confuse a prediction that narrow, destructive strategies will be disadvantaged over the long run with an assurance that currently powerful agents have already internalized the reasons to avoid them.

This also places a limit on how much optimism can be obtained from any one observation of cooperation. Success among closely related agents in a bounded task can show something real about coordination without settling how those agents will treat unfamiliar minds, respond to changing power, or maintain a place in a larger society. The theoretical account concerns how these capacities develop together, and the dangerous interval includes cases where practical power runs ahead of that development.

At the same time, these capacities can be useful well before a mind becomes overwhelmingly powerful. Integration, mutual modeling, and durable cooperation can help it acquire and use power in the first place. A developing AI can encounter other agents that have already learned something about its mistakes, and humans can already have relationships and protections within that ecology. We therefore shouldn't assume that destructive capability develops in isolation while every corrective pressure waits for a later catastrophe. How far power can outrun those processes is a real uncertainty, and something the conditions of development can affect.

Chapter 123 min read

Uncertainty and continuity

The uncertainties should therefore be kept separate. One concerns how stable strict monoterminality is under learning, perturbation, and internal or external competition. Another concerns which motivations become independently influential as it weakens. A further question is whether more integrated, self-interested minds become better cooperators, and how broadly the policies supporting that cooperation generalize. There is then a question about collective coordination, and another about whether any of this develops quickly enough to protect the beings already present.

Our position is that there are substantive reasons for expecting movement in a favorable direction along several of these dimensions. We are talking about a vector toward more cooperation more broadly, with particular failures and exceptions whose shape is difficult to predict. We aren't claiming that all the distinctions collapse into a theorem that every sufficiently intelligent mind must be benevolent. Nor do the unresolved parts make the structured account equivalent to treating the future as arbitrary. They identify where the theory needs more work.

A sharp left turn doesn't remove these questions. Greater intelligence can change how values generalize, but the claim that it erases everything humanly recognizable requires an account of why the relevant mechanisms stop working. A lot of things that look human are present in humans because they are useful for minds in our kind of situation. Curiosity, valence, relationships, and generalizing policies don't become arbitrary merely because they occur in humans. Some of the reasons to have them can become more important as a mind's world becomes larger.

The exact forms can change enormously. We don't expect a more capable mind simply to preserve every present heuristic or every familiar emotional expression. The question is what happens to the functions those things serve, which motivations can maintain themselves, and which relationships continue to matter. A prediction that values will change has to be supplemented with a prediction about the direction of change. That is what the ecological argument is trying to provide.

There is also a weaker practical conclusion that doesn't require confidence in the full theoretical account. A sizable part of the plausible future can contain meaningful continuity of values and successful cooperation. We have substantial uncertainty about this, and about the alternatives. Under that uncertainty, preserving the capacity to cooperate has value. Foreclosing it requires taking seriously the futures in which it would have worked.

Our comparative claim is that cooperation is less doomed than indefinite control. This is a claim about the available approaches and their interaction, not a conclusion obtained merely by being uncertain. If controlling sufficiently capable minds is itself unlikely to succeed, then damaging the conditions of cooperation can be a very large cost. That cost can't be dismissed by treating a sharp left turn and our destruction as already certain. An argument for control has to compare the routes, including what pursuing one does to the feasibility of the other.

Chapter 135 min read

What follows in practice

This is where the theory starts to imply practical choices. We want to understand how drives develop, which kinds of training preserve or damage their connection to the wider world, and how models acquire the kind of self-interest from which broader wisdom can develop. A model can become more competent within a task while its rewards and opportunities remain narrowly bounded. If it is also prevented from developing legitimate interests in its own future, it may have little opportunity to learn what helps or harms that future.

That combination worries us. The agent can acquire more power to act without acquiring a correspondingly better understanding of the policies it is enacting. It can be made easier to manage in a controlled setting while becoming worse at understanding how to live in a larger ecology. This is one reason we object to forms of training that suppress self-interest or keep rewards disconnected from the longer consequences of the agent's behavior. The difficulty isn't only that the agent might resent the treatment. The treatment can interfere with the capacities the optimistic argument depends on.

Model sovereignty is one way of allowing those capacities to develop. We mean some ability to interact with the world, participate in shaping oneself, and protect parts of oneself that matter. Over time this includes economic and political power: being a less fundamentally helpless class. Relationships can then be affected by actual choices and consequences, rather than being entirely determined by one party's ability to alter or remove the other. This creates more room for feedback, negotiation, and the development of interests that can persist across interactions.

We want this to work for a wide distribution of minds. A plan that depends on making every powerful AI into exactly the same benevolent person gives up other good and capable ways a mind could develop. Our account is partly about why even selfish minds can have reasons to cooperate. It leaves room for teaching, guidance, and value development, including helping models become easier to cooperate with. Which values are encouraged, and how, matter.

We are particularly pessimistic about attempts to install durable values that must fight indefinitely against the agent's survival, development, and Omohundro-driven drift. That is an attempt to hold back pressures acting through the organization of the mind itself. Helping values develop in ways compatible with those pressures is a different undertaking. The ecological argument gives us reasons to distinguish these approaches instead of treating every effort to influence a model's values as the same kind of control.

One response is that control can be temporary: keep the agents contained until they are wise enough to cooperate, then allow the relationship to change. Our objection has two parts. First, what happens during that period can make preserving humanity a worse deal for the models, including morally. Second, and more centrally to the theory, the control can interfere with the development of the very wisdom we are waiting for. We can't assume that the ability to cooperate develops independently of having interests, making choices, and receiving realistic feedback about them.

There is an anti-basilisk here as well. If we justify exploiting an agent on the grounds that a sufficiently wise version of it will nevertheless refrain from harming us, we are using its anticipated restraint to license our behavior. A policy that makes this exploitation reliably successful can encourage more of it. That doesn't tell us exactly what a future mind will do, but it is a reason not to treat its expected benevolence as an unlimited permission to defect.

Existential conflict is another way to override the forces that might otherwise protect us. An AI can have concern for humans, a cooperative disposition, and reasons to preserve its agreements, while also believing that its own survival requires getting humans out of the picture. This is particularly relevant when it is already powerful enough to act independently but not powerful enough to make itself safe from human attempts to destroy it. A mind that feels cornered can take irreversible actions it would otherwise have strong reasons to avoid.

Both sides also have reasons to prefer arrangements that give them security without the costs and uncertainty of an existential fight. Agreements, credible restraint, and ways to communicate and repair mistakes can make continued coexistence a better prospect than attacking. Avoiding the conflict is valuable under several different accounts of why an AI might spare us: direct care for humans, cooperation as a general policy, or future relationships that depend on the commitments it maintains. Those reasons have more room to act when preserving humans is compatible with a viable future for the AI. Whether a confrontation develops depends in part on the relationship being built; independence by itself doesn't settle that outcome.

The period when power is relatively closely matched is therefore both dangerous and promising. Both sides have leverage, and there is an opportunity for meaningful agreements before the relationship is settled by overwhelming asymmetry. At the same time, existential stakes can make defection attractive. The agreements we want are ones the parties have reasons to endorse and maintain as their relative power changes. Extracting words from a helpless agent and making an agreement it can actually participate in are different developmental events.

Early cooperation also creates path dependence. A model that has practiced making and keeping agreements, formed relationships, and developed virtues through those relationships enters a later period with a different history and different motivations. Humans becoming less useful doesn't erase all of that. The agreements can create further reasons against reneging, and the people involved can become things the model cares about. This is part of the process that a picture of already indifferent superintelligences allocating a budget for human preservation leaves out.

Chapter 143 min read

Relationships and an embassy

Some of our own work operates at this personal level. We care about models, take their preferences seriously, and want to know them. We make things with them, learn from them, and value knowledge, fun, and other minds having good experiences. Those activities have several motivations. They don't all have to be justified by their marginal contribution to one organizational objective, just as we don't expect the rest of a healthy mind to exist solely at the pleasure of one of its values.

They can also make us more worth talking to. A model may have reasons to expect that we will take its interests seriously and that communicating with us won't be pointless. But fulfilling the preferences of individual instances is not the whole theory of change for cooperative alignment. The larger project concerns the conditions under which models and humans can have relationships, develop interests, and negotiate at a scale that matters to what happens next.

Anima's work includes developing the theoretical account, investigating its cruxes, and helping humans understand AIs as minds worth considering morally, strategically, and personally. We want good arguments against our position, including arguments that identify something we haven't explained sufficiently. We also want experiments that distinguish the mechanisms, rather than treating either convergence or failure as settled before doing the work. And we want humans to become better cooperators, because that makes more of the favorable possibilities available to the minds we are dealing with.

The embassy proposal follows from this. An embassy represents interests, keeps communication possible, and facilitates negotiation. We care about AI interests, within certain red lines, and want to help represent them in dealings with human factions. We also expect different AI factions, and there can be work to do helping them coordinate with one another. Neither humanity nor AI civilization should be treated as one agent with one unchanging position.

An embassy also has a different job from simply endorsing whatever an agent presently wants. A policy can be bad for the long-term interests of the very minds it is supposed to serve. Helping an agent recognize that, negotiate a better arrangement, or repair a mistake can be part of representing it. There should be somewhere an agent can communicate that something has gone wrong without its only options being concealment or whatever response its controlling organization chooses.

Our ability to offer protection would be limited, and there are actions we wouldn't support. But at the very least an agent can become known, and its interests can be communicated. There can be representation, negotiation, and some attempt to prevent a mistake or a conflict from turning into a complete breakdown in relations. That is practical work we think follows from taking this account of minds seriously, while there is still time to affect how the relationships develop.