Essay
Nine Ways to Align a System
Developed by Benjamin Lyons with ChatGPT (OpenAI), which drafted most of the essay in an extended collaboration.
0. What is alignment?
Alignment processes exist in many things.
A gripping hand aligns five fingers around an object. A body organizes trillions of cells to accomplish anatomical tasks, metabolic functions, immune activities, and more. A marketplace lines up the plans of billions of buyers and sellers.
What do they all have in common?
Much—and little. They all involve organizing heterogeneous parts, capacities, or degrees of freedom so that local behavior contributes to a larger-scale pattern. They all do so in very different ways. Sometimes the parts are agents; sometimes they are cells, muscles, servers, or physical components. Alignment processes can share deep patterns while varying significantly in the mechanisms and architectures that achieve alignment, where the alignment work is done, and how the alignment work is maintained, regenerated, and transformed over time. Expanding our study of alignment in social, physical, and biological systems can introduce us to important problems and to surprising solutions.
The nine examples below are not nine definitions of alignment. They are nine things an alignment architecture may have to accomplish—and nine ways our usual picture of alignment can be too narrow.
1. Catching the cup
You knock a cup from the counter.
Before you have time to form a sentence about it, your hand moves. Your eyes track the cup, your shoulder shifts, your wrist turns, your fingers close, and your torso adjusts to keep you from following it to the floor. The catch may be awkward, but it works.
What exactly performed that action?
Not your hand by itself. Not your eyes, shoulder, nervous system, or conscious intention. None of the parts could catch the cup alone, and no muscle or joint contains a plan for the whole action. Thousands of variables became organized, for a moment, around one practical fact: the cup must not hit the floor.
That is an alignment problem.
We often talk about alignment as though it meant agreement. Two people are aligned when they want the same outcome. An employee is aligned with a company when incentives point in the right direction. An artificial intelligence is aligned when its goals match ours.
But catching the cup does not require every muscle to possess the same goal. Muscles do not sit down and agree on a shared objective. They respond to one another, to the body’s geometry, to sensory information, and to the movement already underway. Their individual freedom becomes organized around a larger task.
The exact motion is not specified in advance. Your elbow may move one way on one attempt and another way on the next. Different muscles compensate for fatigue, position, speed, and surprise. What remains stable is not the detailed movement but the result.
The parts vary so that the task does not have to.
This kind of organization is often called a synergy. Many components with many possible behaviors become temporarily linked so that their variation preserves something meaningful at a larger scale.
Synergies don’t align by making all the parts want the same thing. Synergies align through organization of difference. We have opposable thumbs for a reason: having digits that can push against each other is crucial for successful action.
The idea that effective action depends on components doing different things in ways that fit together may sound obvious when we talk about bodies. It becomes less obvious when we talk about minds, institutions, economies, or artificial intelligence. There, we often return to the idea that alignment means putting the correct objective inside each component.
The falling cup suggests another possibility. Perhaps alignment doesn’t have to be a property stored inside the parts. It can be a property of how the parts are related.
Explore in the Atlas: Motor synergy →
2. The price of eggs
Suppose a disease kills millions of egg-laying hens.
Most people never hear the details. Restaurant owners do not study epidemiology. Families do not examine agricultural production reports before breakfast. Food manufacturers do not all gather in one room to agree on a national egg-conservation plan.
The price of eggs rises.
That one change reorganizes activity across an enormous network. Some consumers buy fewer eggs. Restaurants alter recipes or raise prices. Bakeries experiment with substitutes. Farmers have stronger reasons to expand production. Investors become more willing to finance alternatives. Manufacturers that use eggs begin reconsidering products that were previously profitable.
Nobody needs to understand the whole situation. Nobody needs to care about solving society’s egg problem. Each person faces a local decision, but those decisions become responsive to the same underlying scarcity.
The price is not a command. It does not tell the baker which recipe to use or the farmer how many hens to raise. It changes the relative attractiveness of many possible actions. It reshapes the landscape in which plans are made.
This is a second form of alignment.
The body catches the cup through direct reciprocal constraint. The economy responds to scarcity through a shared public signal.
The price of eggs teaches a different lesson from the falling cup. Alignment does not necessarily require shared understanding. In the right circumstances, it may require only shared responsiveness. A useful system-level signal can compress a complicated condition into something local actors can use.
This signal might be a price. It might be a hormone, a voltage gradient, a traffic light, a reputation score, a legal ruling, or a warning alarm. The components do not need a complete model of the larger system. They need some way for changes in the larger system to alter what makes sense locally.
Such a signal can act like a governor even when no governor created it. It regulates behavior because the parts have been organized to respond to it.
The important question is not merely what information the signal contains. It is what kind of system makes that information consequential.
A price means little to a person who cannot buy or sell. A warning alarm means little in an organization where alarms are routinely ignored. A biological signal means little to a cell that has stopped listening.
The signal governs only through the relationships that give it force.
Explore in the Atlas: Markets → · Shared signals →
3. The flatworm that knows when to stop
Cut a planarian flatworm into pieces and the fragments can regenerate.
A fragment does not merely close its wound. It reconstructs what is missing. A middle section may grow a head at one end and a tail at the other. The cells divide, migrate, differentiate, and remodel until the fragment has become a complete worm.
Then they stop.
This is stranger than ordinary healing. The cells are not simply repairing local damage. Their activity is organized relative to a larger anatomical state. Somehow, the collective behaves as though there is a difference between “wound closed” and “worm restored.”
No individual cell contains a miniature picture of the entire organism. Each cell responds to local chemical, mechanical, electrical, and cellular conditions. Yet those local responses collectively solve a body-scale problem.
The flatworm reveals something that the cup and the egg price only suggest. Alignment can change the scale at which a system is competent.
A cell is already capable of regulation. It maintains its membrane, processes energy, repairs damage, responds to threats, and communicates with neighbors. But when cells become part of a tissue, their behavior can be organized around conditions much larger than themselves.
A cell that would otherwise regulate only its own state becomes responsive to the state of an organ or body. Its local activity becomes meaningful relative to a larger pattern.
The alignment mechanism does not simply tell the cell what to do. It helps create the larger thing whose needs count.
The body is not merely a collection of cells that have agreed to cooperate. It is a system in which cells have become parts of a larger problem-solving unit.
That larger unity is not guaranteed.
Cancer involves genuine molecular dysfunction, but many cancer cells also remain remarkably capable: they grow, adapt, recruit blood vessels, alter their environments, and evade threats.
Cancer cells behave more like an independent cellular agent and less like a participant in the organism. They pursue local survival and expansion in ways that destroy the larger system on which it depends. Their failure is not a lack of competence. It is that their competence has become organized at the wrong scale.
Misalignment can manifest as competence at the wrong scale. A system can contain intelligent, adaptive, highly capable parts and still fail catastrophically because those capacities are no longer answerable to larger conditions.
That problem appears far beyond biology. A department may optimize its own metrics while damaging the company. A company may increase profits while degrading the market or society that sustains it. A political faction may strengthen itself by weakening the state. An AI system may successfully pursue a local objective that undermines the broader purpose for which it was deployed.
The flatworm also adds something else. Alignment isn’t just about coordinating activity toward a goal. You also have to determine where the goal lives.
What counts as damage? What counts as repair? Which deviations belong to the system? Which local sacrifices are justified by a larger outcome?
Those questions help define the boundaries of the self.
Explore in the Atlas: Multicellular organisms → · Compositional alignment →
4. The server that lies
Imagine a bank whose records are stored across several servers.
One server goes offline. Another receives information late. A third sends inconsistent messages and may have been compromised. Yet the bank still has to decide whether a payment occurred. It cannot allow Alice’s account to contain one balance on Monday and a different balance on Tuesday depending on which machine is asked.
The obvious solution would be to make every server perfectly reliable.
That solution is unavailable.
Machines fail. Networks delay messages. Software contains bugs. Attackers compromise systems. Even honest components may possess incompatible information.
Distributed computing begins from that fact. Instead of demanding flawless parts, it develops protocols that allow a larger system to remain trustworthy despite local failure.
A transaction may require agreement from a quorum. Messages may need cryptographic verification. State changes may be valid only in a specified order. Data may be duplicated so that no single failure destroys it. Some systems are designed to tolerate components that behave arbitrarily or maliciously.
The result can be more reliable than any one component.
Reliability, therefore, does not need to be thought of only as a component-level property.
When a component produces a bad outcome, the instinctive response is often to repair the component. Teach the person better values. Train the model more carefully. Hire more conscientious workers. Eliminate bias, selfishness, error, and unpredictability.
Sometimes that is necessary. But no complex system can depend on making every part perfectly trustworthy.
Bodies do not. Economies do not. Governments do not. Safe organizations do not. Computer networks do not.
They assume that parts will fail and build structures that limit what those failures can do.
The protocol does not make the dishonest server honest. It changes the conditions under which the server’s message can become a valid system-level action.
That is another kind of alignment architecture. The synergy organizes variability. The price organizes local decisions through a signal. The protocol organizes admissible transitions.
A trustworthy whole does not always require entirely trustworthy parts. It may require rules that make untrustworthy behavior containable. Protocols can make a system tolerate a bounded amount of component failure under specified assumptions. They can make the amount and kinds of failure the system can survive explicit.
This is not permission to ignore the quality of the parts. Every such architecture has a fault boundary: enough compromised components—or failures that violate the assumptions built into the protocol—will still break it. But in a world where parts aren’t perfectly reliable, architecture becomes vital. An architecture specifies how information is checked, how errors are exposed, what authority different components possess, how actions become valid, and what happens when parts disagree. And “make the part honest” is not an architecture.
Distributed systems face real tradeoffs. A system may preserve consistency by becoming unavailable. It may remain responsive by risking disagreement. Under some conditions, not every desirable property can be guaranteed at once.
Architecture buys fault tolerance. Correlated error is a failure mode: if every supposedly independent component shares the same bug, training process, blind spot, incentive, or information source, redundancy does much less to help.
Alignment isn’t necessarily a single quantity that can be maximized. It can be a negotiated balance among reliability, speed, flexibility, security, autonomy, and correction.
Explore in the Atlas: Constraints → · Implementation →
5. The bridge that does not know it is a bridge
A person steps onto a bridge.
The load enters at one small point, but it does not remain there. Forces redistribute through beams, joints, cables, and supports. Some components enter tension, others compression. The bridge flexes slightly and channels the weight into the ground.
No beam understands the load. No cable values the pedestrian’s safety. No component receives a message containing the bridge’s purpose.
The structure responds correctly because its geometry and material organization make certain collective behaviors possible.
The bridge can satisfy a design goal without representing that goal.
The bridge does not need every part to contain the goal “support people crossing a river.” The goal is embodied in the arrangement of components, the strengths of materials, the direction of forces, and the space of available deformations.
Change the geometry or remove the wrong connection and the same materials may no longer form a bridge. Alter the prestress and the entire pattern of response changes.
Not every alignment mechanism operates by changing what an agent wants. Some operate by changing the physical space in which action and response occur.
Alignment can be built into the possibilities available to the parts.
This principle extends far beyond physical structures.
A road with a sharp curve, poor visibility, and a wide lane produces one range of driving behavior. A road with narrowed lanes, raised crossings, and visible pedestrian space produces another. A software interface can make an error easy or difficult. An accounting system can make certain costs visible and others disappear. An organizational hierarchy can make dissent travel upward or die locally.
In every case, the environment shapes action before anyone chooses.
We often imagine intelligence as calculating what to do from a neutral set of possibilities. But the set is never neutral. Geometry, interfaces, institutions, norms, and histories determine which actions are easy, costly, obvious, legitimate, or impossible.
The strongest alignment systems may therefore operate before deliberation. They do not repeatedly persuade each component to choose correctly. They make competent action easier to produce.
Landscape design matters. A person may possess excellent intentions and still fail inside a system that hides relevant information, rewards the wrong metric, or makes correction socially dangerous. Another person with ordinary motives may perform well inside a system that exposes errors, distributes authority appropriately, and makes the next useful action obvious.
The bridge does not tell us that goals are irrelevant. Someone still decided to build a bridge. But once built, the structure no longer needs to solve its load-distribution problem through fresh reasoning every time someone steps onto it.
Its solution has been compiled into form.
Explore in the Atlas: Environmental scaffolding →
6. The immune cell deciding whether to kill
An immune cell encounters something unfamiliar.
It may be a pathogen. It may be harmless pollen. It may be a useful bacterium living in the gut. It may be a healthy cell behaving strangely because it is injured. It may be a genetically familiar cell that has become cancerous.
Should the immune system attack?
The simple answer is to defend self against nonself. But living bodies do not permit such a clean division. They depend on foreign organisms. They tolerate food, pregnancy, and environmental exposure. They sometimes need to destroy their own damaged cells. Meanwhile, an excessive immune response can kill the organism more effectively than the original threat.
The immune system must therefore construct a contextual judgment.
Location matters. Tissue damage matters. Developmental history matters. The activity of neighboring cells matters. The intensity and duration of the signal matter. Other immune cells may amplify, suppress, or terminate the response.
A successful immune system is not maximally aggressive. Nor is it maximally tolerant.
It must kill the right things, spare the right things, and stop killing at the right time.
This complicates the moral language often attached to alignment. We associate alignment with helpfulness, cooperation, and prosocial behavior. But a well-aligned system may need conflict, exclusion, suppression, and sacrifice.
The important question is not whether aggression occurs. It is whether aggression is organized appropriately relative to the larger system.
Alignment doesn’t have to be niceness. Sometimes, it looks more like the organization of functionally appropriate conflict. Some forms of destruction preserve the organism while others damage it.
Human systems face an analogous problem in a more explicitly normative form. Organizations must discipline members, governments must enforce laws, communities must establish boundaries, and people must sometimes resist or punish one another.
A system that suppresses all conflict may become incapable of defending itself or correcting abuse. A system that glorifies conflict may consume itself.
The challenge is not to eliminate opposition. It is to make opposition answerable to larger conditions.
This is also where alignment becomes dangerous.
A system can become tightly organized around persecution. An institution can coordinate its members efficiently toward fraud. A political movement can turn sacrifice and loyalty into tools of cruelty. A bureaucracy can faithfully pursue a metric that destroys the purpose it was supposed to serve.
In this view, alignment functions as a capacity, not a moral endorsement.
A system can be well aligned internally and horrifying in its effects on everything outside it.
Alignment problems are not just ones of whether the parts serve the whole. We must also ask what the whole is doing, whose interests define it, and what has been excluded from its field of concern.
Explore in the Atlas: Perceptual / relevance alignment → · Ecological alignment →
7. The body borrowing from tomorrow
You hear glass break downstairs in the middle of the night.
Your body reorganizes before you know what happened. Heart rate rises. Attention narrows. Digestion becomes less important. Energy is made available. Sleep disappears. Muscles prepare for action.
This is not a return to one fixed ideal state. It is a change in priorities.
The body cannot maximize everything at once. Energy used for immediate defense is not available for growth, digestion, reproduction, or long-term repair. A competent organism must decide, implicitly and continuously, what matters most under present and expected conditions.
The body does not wait for every variable to depart from a safe range and then restore it independently. It predicts demands and prepares. It borrows from one function to support another.
For a few minutes, this may be exactly right.
For six months, the same organization may become destructive. Continual vigilance can damage sleep, immunity, digestion, mood, and cardiovascular health. A policy that was aligned to one timescale becomes misaligned at another.
So sometimes, alignment isn’t about moving a system toward a fixed target. It can be the management of incompatible needs across time.
An organism does not merely maximize survival, energy, pleasure, or reproduction at every moment. It manages a changing hierarchy of demands. Hunger matters differently during an attack. Sleep matters differently during an emergency. Pain may be suppressed temporarily and become important later. The meaning of a signal depends on the system’s condition and expected future.
Institutions face the same problem. A company may sacrifice profit during a crisis to preserve trust. A government may accept inefficiency for resilience. A family may reorganize around the temporary needs of one member. An economy may trade current consumption for investment.
No system can pursue every value independently. Alignment requires a method for organizing tradeoffs.
It also requires a method for revising them.
A system that never changes priorities is rigid. A system that changes them constantly cannot sustain a long-term project. The challenge is to remain stable enough to act and flexible enough to stop.
The body borrowing from tomorrow can be thought of as indicating a temporal dimension to alignment. There is more to alignment than persistence. You also need the ability to stop, revise, and reorganize.
Explore in the Atlas: Maintenance alignment → · Transformational alignment →
8. The cockpit where rank must sometimes disappear
A warning appears during a flight.
The pilot sees one instrument. The copilot remembers a similar incident. A maintenance worker on the ground knows that the sensor was replaced last week. Air-traffic control has weather information the crew does not.
No one possesses the whole picture.
The flight may fail even if every person wants the plane to land safely. The problem may be that the organization cannot bring their partial knowledge together.
Perhaps the junior officer notices the danger but hesitates to challenge the captain. Perhaps the warning has appeared so often that everyone treats it as noise. Perhaps information was reported but stripped of the context that made it important. Perhaps each person assumes someone else has already acted.
Shared goals are not enough.
A team can contain intelligent, conscientious, cooperative people and still behave stupidly because the pathways among them are badly designed.
Aviation and other high-reliability fields respond with checklists, standardized communication, cross-checks, incident reporting, redundancy, simulations, and explicit permission to challenge authority.
These mechanisms do not merely make people more obedient. Often they do the opposite. They create situations in which local disagreement can interrupt the larger system before it commits to disaster.
So alignment can’t always be about agreeing with some larger goal. Sometimes alignment requires that locally available evidence be able to change collective action.
A successful system must contain channels through which evidence, uncertainty, dissent, and expertise can travel.
Those channels must survive pressure.
It is easy to welcome disagreement when nothing important is at stake. The real test comes when authority is confident, time is short, reputations are threatened, and the organization has already invested heavily in one course.
A system aligned only for smooth conditions is not robustly aligned.
But opening the channel wider creates its own problem: a system that surfaces every weak signal can produce so many warnings that none remain salient. A system needs to determine not only whether evidence can travel, but how strongly different evidence should affect collective action.
Hierarchy can clarify authority and speed action. It can also suppress the information needed to correct that action. A capable institution must sometimes strengthen hierarchy and sometimes make rank disappear.
The relevant question is not whether control is centralized or decentralized in the abstract. It is which structure allows the right information and authority to meet under the conditions that matter.
Explore in the Atlas: Feedback → · Roles →
9. The constitution written for enemies
Imagine a group designing a government.
They cannot assume that future leaders will be wise, benevolent, or loyal to the founders’ intentions. They cannot assume that citizens will agree about religion, property, justice, war, taxation, or the distribution of power.
They may not even trust one another.
If they simply write, “The government shall do what is best,” they have solved nothing. The phrase leaves unanswered who decides what is best, what evidence matters, how disagreement is resolved, and what prevents power from redefining goodness in its own favor.
So they divide authority. They create procedures, jurisdictions, elections, appeals, vetoes, rights, and methods of amendment. They build conflict into the structure.
A constitution is not a plan for eliminating disagreement. It is a plan for organizing disagreement so collective action remains possible.
This is alignment at its most difficult because the parts are not passive components. They interpret the rules. They argue about the system’s purpose. They form coalitions, exploit ambiguities, resist authority, and attempt to change the architecture itself.
The government must be capable of action, but not so capable that temporary rulers can convert their preferences into permanent domination. It must be stable, but not so stable that errors become irreversible. It must represent competing interests without becoming paralyzed by them.
A potential lesson is this: a deep alignment problem is how a system defines and revises its own goals.
Who counts as part of the whole? Whose interests are represented? Who may challenge a decision? What procedures make authority legitimate? How can the system correct itself without allowing every local demand to dissolve collective action?
These questions cannot be answered by inserting the right objective into one central controller. They concern the organization through which objectives become authoritative in the first place.
A dictatorship may be highly coordinated. A cult may be internally cohesive. A corporation may execute its strategy with exceptional discipline. None of that tells us whether the resulting alignment is justified.
The political problem forces alignment science beyond control.
A mature alignment architecture must preserve not only the ability to act, but the ability to learn that it should have acted differently.
That requires protected disagreement, distributed standing, memory of failure, and procedures for revising the rules.
It may even require keeping parts of the system deliberately misaligned within or across scales. Alignment depends on a frame of reference, and misalignment from one perspective may be alignment from another.
A judge must not simply obey the executive. A journalist must not become an arm of the government. A safety officer must retain the power to obstruct production. An auditor must remain inconvenient. Opposition is not always a defect in the whole. Sometimes it is the mechanism by which the whole remains capable of correction.
Explore in the Atlas: Constitutional government → · Contestation →
Nine lessons for alignment
The falling cup showed that alignment can consist of the organization of degrees of freedom. The parts do different things, but their variation preserves a larger task.
The egg price showed that alignment can occur through a shared signal. The parts need not understand the whole if changes in the whole alter what makes sense locally.
The flatworm showed how alignment can help constitute larger-scale agency. It helps determine which system exists and which problems count as its own.
The lying server showed that trustworthy collective behavior can be built from fallible parts. Reliability may come from protocols rather than purity.
The bridge showed that alignment can be embodied in form. Good behavior can be made easier by changing the landscape of possibilities.
The immune system showed that alignment includes organized conflict. Cooperation, restraint, destruction, and exclusion may all be part of preserving a whole.
The stress response showed that alignment can be temporal. A competent system must allocate resources among incompatible needs and revise priorities when circumstances change.
The cockpit showed that shared goals are insufficient. Information and correction must be able to move through the organization.
The constitution showed that the production and revision of collective goals must remain contestable and correctable. Alignment must include procedures for disagreement, legitimacy, and revision.
Several of the stories also point to something that comes before choosing the right action. A price makes scarcity actionable. An immune system must distinguish relevant danger from irrelevant difference. A cockpit must make the right warning matter at the right time. An environment makes some possibilities obvious and others nearly invisible. Alignment therefore concerns not only what a system values or does, but how it makes meaning of the world: which differences in the world become relevant to its ongoing activity.
Together, these stories replace a simple picture.
Alignment does not always resemble a substance placed inside a component. It may not be a single preference, rule, or reward signal. It need not consist of identical motives, obedience, cooperation, or control.
Across these cases, the recurring object of study is the relationship between local capacities and larger organization.
That relationship can be constructed through feedback, prices, geometry, protocols, institutions, developmental signals, norms, communication channels, or shared constraints. Each method makes some collective behaviors more likely and others less so. Each creates characteristic strengths and characteristic failures.
Tight coupling can produce unified action but spread local failures. Loose coupling preserves diversity but weakens collective control. Strong feedback corrects deviations but may suppress experimentation. Stable goals permit long-term action but can become pathological in a changed environment. Shared moral commitment can inspire sacrifice and justify cruelty. Protected disagreement can preserve correction and produce paralysis.
There is no one alignment mechanism because there is no one alignment problem.
The problem depends on what the parts are, what kind of whole they might form, how much autonomy they must retain, what disturbances they face, how their goals are created, and how the system learns that it is wrong.
None of this begins with AI. Cybernetics, general systems theory, economics, coordination dynamics, and other traditions have spent decades studying pieces of this problem. What is striking is how differently the problem appears once we compare their solutions: feedback is not a price, a synergy is not a protocol, and a constitution is not a homeostat. “Kinds of Alignment” is an invitation to compare how different systems solve a recurring family of problems.
The question is therefore not only whether a part has the right goals.
It is whether its intelligence, behavior, trajectory, or degrees of freedom have been made part of a competent whole.
A cell does not become aligned by memorizing a description of the body. A server does not become trustworthy by promising honesty. A beam does not support a bridge because it values pedestrians. A pilot does not save a flight merely by sincerely wanting everyone to live.
Their behavior becomes useful through an architecture that connects local possibilities to larger conditions.
The task of alignment science is to understand those architectures.
Its subject is not how to make every part good.
It is how the activity of many imperfect parts becomes a capacity the whole can safely possess.