Essay

Kinds of Alignment


1. Introduction: There are many kinds of alignment

Consider a few different examples of aligned systems. A market coordinates people who want different things—buyers and sellers all interacting without anyone agreeing on what the optimal consumption bundle is or what the overall economy ought to be like. A soccer team has a shared goal, unlike a market, but all the players must perform different roles to collectively achieve that goal—the goal itself doesn’t specify anything about how a soccer team should be organized or who should play at what position. A firm is highly alignable: it’s very responsive to profits, prices, contracts, regulation or the threat thereof, etc., without possessing the values of whoever steers it—a firm that responds as intended to a Pigouvian tax may not care about the social goal the tax is designed to achieve. And collective action problems consist of many individual people each behaving in a perfectly sensible way whose collective behavior could be seen as inconsistent with some social goal.

All of these systems involve alignment, but the kinds of alignment are not the same. A market exhibits a high degree of coordination with little common agreement about individual or collective goals. A soccer team has a high degree of agreement about the shared goal, but achieving that shared goal requires deliberate differentiation into different roles, each of which handles a local set of problems rather than directly pursuing the goal of maximizing win probability. A firm exhibits a high degree of responsiveness to external signals, making its behavior very alignable, but it does not share the values of its external governors. And collective action problems show how individually sensible behavior can produce something collectively undesirable, with no apparent fault as to the individual goals but rather in how they compose.

Properties that all feel like “alignment” in one context or another do not necessarily go together. They can come apart—coordination, value agreement, centralized behavior, alignability, acceptable local objectives—none of these things necessarily imply the other. So alignment may not be a single property varying from high to low. I propose it is a family of distinguishable organizational relations that can come apart across systems. “Kinds of Alignment” isn’t about asking whose values should the AI share. It asks, “What phenomena have we been grouping together under ‘alignment’?”

2. From agreement to concordance

The most intuitive way to think about alignment is sameness of values. In this view, alignment occurs when the components share the goal of the higher-scale system, like AI agents sharing “human values”. But many systems exhibit a very different kind of alignment: alignment as concordance, a compatibility of actions, plans, or trajectories. Markets are an example: prices bring people’s production and consumption plans into mutual compatibility, even if everyone is totally self-interested.

Concordance and agreement don’t always go together. Markets exhibit concordance without agreement: economic agents pursue different ends while prices nevertheless make their plans compatible. But a poorly managed team of workers may exhibit agreement without concordance: each worker can genuinely try to accomplish the firm’s mission and nevertheless interfere with each other.

Concordance may even involve the intentional, deliberate crafting of differences. A soccer team does not become more aligned when the goalkeeper behaves more like a striker. And rather than having each team member think about the full-scale goal of winning the match and trying to execute an appropriate strategy, the team benefits by having each player stick to their own domain of activities and play their intended role in a larger structure. The goalie tries to do their job as a goalie; they do not try to “win the match.” And if the goalie’s ultimate objective is “show I’m the best goalie in the world” rather than “win the match”, they may still be a great goalie—possibly better!

Examples like markets and teams motivate a conceptual pivot. These are recognizable forms of alignment, but they don’t require or even necessarily benefit from having each agent pursue a shared system-level value. Instead, these kinds of alignment concern relations among trajectories in coupled systems, rather than similarity among internal states. Market participants couple to each other with budgets and prices; soccer players recognize the developing state of the pitch and perform their role.

Shared values or goals may help produce alignment. But they do not define alignment, are not sufficient for alignment, and may not even be necessary for it.

3. Values, behavior, and alignability: three things we routinely collapse

Think about the stylized profit-maximizing firm. A firm’s behavior typically supplies things that humans value: goods and services that make our lives better. But a firm’s values are not aligned with human values: a firm wants to maximize profit, and, in principle, doesn’t care whether profit correlates with things humans value. Yet the behavior of firms often aligns with human values because firms are very alignable: their behavior can reliably be redirected, corrected, or constrained through prices. If you want firms to sell more pizza, offer a higher price for pizza.

So values, behavior, and alignability come apart. They are not the same thing but can correlate more or less depending on the situation. They’re not all part of one alignment scalar. They are distinguishable dimensions or relations that can vary independently.

  • Value/goal alignment: relevant internal evaluations or objectives agree.
  • Behavioral alignment: behavior reliably fits the relevant larger activity.
  • Concordance: multiple trajectories/plans are mutually compatible.
  • Alignability: behavior can reliably be redirected, corrected, or constrained.

Markets align firms not by teaching them the right values but by relying on their “wrong” values being highly controllable through prices. This means that alignment can be achieved through architecture. That points to another question: how is the alignment produced?

  • Architecturally produced alignment: trajectories are shapeable through incentives, constraints, protocols, feedback, institutions, coupling, etc.

It’s also striking that firms having the “wrong” values is precisely what makes them so alignable. If firms cared about social goals in addition to profits, they might sometimes neglect profit signals. Chick-fil-A illustrates what happens when the idealized profit-maximizing model weakens: it doesn’t open on Sundays even though they could make money by doing so because of their commitment to religious values.

No single type of alignment is ideal in isolation. Observing good behavior doesn’t tell us why the system behaves well. It could be because the system’s components have good values, or because they are all very alignable, or because the alignment architecture is good, etc. Nor does possessing good values tell us whether the component can be aligned successfully. Bad signals, poorly designed constraints, or a lack of relational coupling can make a good component behave badly with respect to the system’s goals.

4. Alignment changes with scale: aligned parts do not imply an aligned whole

Aligned parts do not imply an aligned whole. Even if parts have good values or are behaving in ways aligned with the larger system, this may not compose across scales. Someone who is deeply dedicated to their family may be willing to lie, cheat, or steal for them. A profit-maximizing firm may discover that a production process that creates a negative externality is the best way to make money. A competitive athlete may break the rules. And so on.

This makes it very hard to determine if a component is aligned. People should care about their families. Firms that try to maximize profits are important to the economy. Sports wouldn’t be as entertaining if the athletes didn’t obsess over winning. In some circumstances, a component will have the right values or behavior; in others, it won’t.

Call this kind of alignment “compositional alignment.”

  • Compositional alignment: does an alignment relation present at one scale survive construction of a higher-scale system?

This gets particularly weird when we think about how goals scale. Trillions of cells behave in ways that create and sustain an individual human being, but a human doesn’t necessarily know what cells are, let alone care about anything a cell cares about. Each cell is working to support a collective pattern that involves goals that are meaningless at the level of the individual cell.

The economy exhibits something similar. No individual firm or person needs to know what the optimal allocation of resources is, let alone care to achieve it. Higher-scale goals can be qualitatively transformed relative to component goals. This transformation helps keep things aligned, constantly discovering what large-scale pattern makes everyone’s plans mutually compatible. Call this “transformational alignment.”

  • Transformational alignment: the collective pattern changes so heterogeneous parts remain coordinated.

As a result, “Is the AI aligned” is underspecified. We need to think about the scale at which alignment is being measured. Alignment at one scale can be misalignment at another, and vice versa—two competing firms may be misaligned with respect to each other, thereby aligning both to our goals of better products at lower prices.

5. A multidimensional framework

Because there are many kinds of alignment, it doesn’t make sense to say, “Is this agent aligned?” or even “Is this system aligned?” Different kinds of alignment are neither additive components of a single alignment scalar nor mutually exclusive. They answer different questions about the same system. A system should thus be characterized by an alignment profile, not assigned one alignment score or one alignment type. Table 1 depicts some dimensions a profile might consider.

Dimension Question Examples
Target Aligned to what or whom? user, principal, institution, collective condition, environment
Manifestation What is aligned? values, goals, perceptions, models, plans, behavior, error correction
Mechanism How is alignment produced? agreement, incentives, constraints, feedback, protocols, coupling
Scale Where does it hold? component, dyad, organization, society, organism
Time What does alignment do under change? achievement, maintenance, correction, regeneration, development, transformation, generation
Tradeoff / quality What does it preserve or sacrifice? autonomy, diversity, repertoire, robustness, corrigibility

Table 1: Dimensions of an alignment profile

Kinds of alignment do not form a flat taxonomy. Alignment can differ along several analytically distinguishable dimensions that can come apart and interact—what is aligned, to what, by what mechanism, at what scale, and with what dynamics over time. Familiar ‘kinds’ of alignment pick out recurring relations or profiles within this multidimensional space. For example, a simple profile for a market system would be: low shared goals, low shared values, high concordance, strong signal-mediated architecture, broad scale, substantial externalities, considerable adaptive capacity.

6. Three consequences for alignment research

Disentangling different notions of alignment has several important consequences for alignment research.

First, more alignment is not necessarily better. A group can exhibit unusually high agreement in values and goals while pursuing disastrous ends. Excessive alignment can also produce rigidity, groupthink, lost autonomy, suppressed disagreement, and reduced repertoire. Dialing up a particular dimension of alignment can make the system work worse, not better.

Second, there are genuine reversals that can occur as you improve along one dimension of alignment. As you improve in epistemic similarity, the collective can have worse accuracy—individual disagreement is helpful for accurate consensus. Going the other way, more behavioral differentiation can lead to better functional alignment, as exhibited by the economy’s division of labor and the positions on a soccer team. The dimensions of alignment aren’t simply imperfectly correlated; they can trade off against each other, and sometimes move oppositely.

Third, “alignment to whom?” is not sufficient. There is more to alignment than a relation between an agent and a target. Market coordination and making the right passes in a soccer game concern relations among interacting trajectories. Instead of just “alignment to whom?” the specification problem becomes something like, “What needs to be aligned, with what, at which scale, through which mechanism, over what conditions and timescale, and with what acceptable tradeoffs?”

Nor is it clear that the specification decision should be made at one point. In the economy, for example, the specification is a collective compromise and shared question-and-answer among billions of interacting agents. The conditions the system is organizing around, along all the various dimensions, are a negotiation, a discovery, and a continual updating.

7. Conclusion: from alignment as a goal to alignment as a science

“Kinds of alignment” become clear when you realize that they separate from each other. Agents with the right values may not exhibit the right behaviors; agents with the right behaviors may not be alignable to new circumstances. Individual agents with the right values and the right behaviors may not compose into a collective with the right values and the right behaviors. There may be no way to specify the right values and the right behaviors at the outset; the collective system formed by the agents may embody a discovery of those things. Misalignment may even be useful; competition within a scale or error signals across scales may be necessary for the system to work well. There can be no single undifferentiated alignment variable when forms of alignment can come apart, and sometimes move oppositely, even when ordinary language conflates them.

This essay has only touched on kinds of alignment. Some of the most important ones I haven’t mentioned already include perceptual/relevance alignment, regenerative alignment, alignment to a virtual governor, constitutive alignment, alignment to a synergy, and strong and weak alignment. But the goal isn’t just a taxonomy. What are the relevant dimensions for kinds of alignment? How do we develop measurable signatures? What are the dependencies and tradeoffs between them? How does alignment change across scale and time? Which kinds of alignment architectures produce which profiles?

There are also engineering implications. Consider the economy as an example. It would be really hard to specify what values each agent in the economy should internalize. But through price coordination, the value requirement becomes astonishingly thin: mainly, it’s just sensitivity to prices. Then the optimizing consumers and producers in the economy discover how to align themselves with each other. Generalizing, we can ask how particular architectures can reliably produce the forms we want—how to compile alignment to translate alignment problems from where they’re harder to solve to where they can be solved more easily.

Keep exploring

Explore the Alignment Atlas →