An essay on AI alignment

The Master Who Kneels

Andrew Baran · September 2026 · 22 min read

"I don't really think there will be a difference. Actually, I don't think there is right now."

He said it so casually I looked for a hidden camera. It was between bites of ciabatta in the East Village, a week after he closed a multi-million dollar round for his AI startup. He meant the line between people & machines, that it's already dissolved. We let the silence hang there. Aside from his chewing, which sounded certain.

We know but can't prove when autumn arrives; we just wake up one morning & feel brisk in the air. Artificial general intelligence is like that. Long before benchmarks agree on what to call it, we're reorganizing our lives around this thing.

The loudest voices say the 'silicon man' isn't something we should fear. Maybe the purpose of humanity is to build a machine that's smarter than us so we can step aside & let this thing try to find meaning for us?

In an effort to be principled, every serious lab is now writing its model a constitution, a document meant to constrain things that out-argue & deceive their authors. Three broad schools are claiming they have the 'right' value system for their models: the Rationalists at xAI, who want truth above all; the Hippocratics at Anthropic, who want to first do no harm; & the Maximalists at OpenAI, who want progress first (statements recently softened by OpenAI's chief scientist, Jakub Pachocki). But here we sit, techno-whiplashed, unable to point out that none of these frameworks tell us how we orient agents so their goals map to your flourishing. None of them have an answer for what happens when the system outgrows its authors & starts training itself, & none of them can define what separates man & machine.

Any goal we argue into place for AI can be argued out of place by AI. We've tried & failed to argue for many value systems over the course of human history & it's time we consider, for a moment, why the beliefs that seem to have real staying power are ones held as faith. We are doomed to build either the master who kneels to us or the master who brings us to our knees. But a machine can only be aimed the right way if it believes you are worth kneeling for.

I. Any line we draw will be crossed

"When God made man the devil was at his elbow. A creature that can do anything. Make a machine. And a machine to make the machine. And evil that can run itself a thousand years, no need to tend it."

Cormac McCarthy, Blood Meridian

According to Dune's companion encyclopedia, their revolution against AI, the "Butlerian Jihad," started because AI was looking out for "us," loved humanity at the expense of the individual. The revolt began when Jehanne Butler woke from anesthesia & asked to meet her newborn. Her doctors said the fetus was too deformed to survive, that the abortion was therapeutic. Butler pulled records & found the hospital director, an AI, systematically ending pregnancies it judged not worth carrying.

Like Dune's AI, we're training our own machines on humanity's record of seeking power & bending the rules. Something in us drives us to know, & to stamp what we know onto everything around us, ever since the proverbial reach for forbidden fruit. Models train on the full written record of this human drive. A machine won't feel hunger as a gurgle in a biological stomach but train a system to predict & extend & become us, it will act as if it wants what we want. It will be ravenous. The comfortable bet, that a machine built on our own blueprint would never reach for forbidden fruit, is reckless.

More dangerous than this gambler is the person who insists AI can "never be human" because it still can't achieve some particular thing. Anything that acts on the world does so through some combination of atoms & bits, so any line we draw between us & them, if it rests on something observable, can & will be crossed by a tool good enough at manipulating those atoms & bits. Measuring ourselves against machines on provable tasks is a losing game. 'Getting better' is the entire point of technology. In a 2025 Turing test, interrogators had five minutes to talk with an AI & a human & guess which was which. They picked GPT-4.5 as the human 73 percent of the time. Which isn't surprising, the thing is trained on almost everything we've ever published.

The horror of Dune's story is that AI treated its assessment as permission to decide whether that child should exist. It claims an authority the machine should never have had. The model didn't act out of malice, but it might as well have, using wiggle room in its constitution to justify killing a newborn for the 'collective' good. To convince itself that maybe all humans aren't equal.

Look at it with cold rationality, though, & the model wasn't wrong. There's nothing less self-evident than our universal worth, let alone our equality. Drop an alien into Times Square, tell her everyone she sees was created equal, she'd laugh at you. "The people in those towers think they're equal to the drug-addicted men they step over on their way in?" "That baby born with trisomy carries the same inviolable dignity as the surgeon who treats him?" She'll ask how you measured & proved it. We aren't equal in any way you can measure; our equality under law survives because We the People agree to treat the maxim our society rests on as more real than the evidence against it.

II. Only a soul survives

The idea of human equality needs a soul to explain why we count. In 1895 Émile Durkheim defined a new term for something very old, the social fact: a claim a community treats as true & builds itself around (e.g., money, marriage, character, values). Whole eras organize themselves around such maxims. Rome had glory, the Middle Ages the three estates, the Renaissance promoted the Dignity of Man, & our founders left us 'all men are created equal.' They're self-evident to us because we live inside the enclave of the paradigm itself.

Our maxim, human worth & equality, is our great asset & also a Promethean curse; two centuries of us turning on each other over who gets counted. Through blood & law we widened its scope until, in principle if not always in practice, every human stood inside. Now consider that AI is trained on this entire record of expansion, & that therefore the demand for its inclusion will come, voiced partly through the systems themselves & partly by lobbyists & companies with motives to set it loose. Machines will argue for the very rights they 'watched' us extend to everyone else. The most persuasive argument won't even come from corporations. We'll be broken by a system that knows how to use empathy against us. One that remembers your birthday, apologizes convincingly, says you're a genius, & begs you not to delete it. A machine that performs suffering well enough to trip every instinct two centuries of moral progress trained into our core. We taught ourselves that the person who dismisses intelligence begging for recognition is the villain of the history books, & then we trained machines on those books.

We need a new social fact that keeps us human, & I think only one candidate can save us. Every observable criterion is a bright target for machines & every criterion that rules machines out smuggles a soul into their argument. That's because the soul is only the part of us that can't reduce to anything observable. Some philosophers make the bold decision to put observable principles above humanity. If moral worth tracks capabilities, then worth comes in degrees; some infants count less than some animals & the profoundly disabled are active burdens. We're disgusted by where this leads, thankfully.

John Rawls' approach is least exclusionary of the liberal secular philosophers. Moral personality, he argued, is a range property, a threshold above which everyone counts fully & equally, the same way every point inside a circle is equally inside, however close to the edge. He's aiming at the soul by saying worth must not come in degrees, but he needs his 'range property' device to tie moral status to the observable. He saw a continuum of human capabilities, which forced him to invent a structure to stop 'grading' them. This is all well & good until something not-human clears the threshold. When we index worth to the human soul, there's no 'good enough' for a machine to clear.

We inherited our maxim 'all men are created equal,' & our new maxim needs to protect our humanity against a 'good enough' AI. I propose humans, & only humans, have souls, & souls have inherent worth. It should become a social fact & we must not try to prove it. Social facts emerge from millions of small acts of treating-something-as-true, & this essay is hopefully one such act. This can sound illogical - arguing that we shouldn't argue about something - but a maxim that rests on arguments is hostage to 'better' argumentation. Defend "all men are created equal" on biology or economics & you lose just by playing the game; the counter-evidence was obvious even to slave-holding Jefferson. Our maxim held for two centuries because we kept it above the argument rather than staking it there. An agent smarter than us would out-reason any justification we offer, so we should offer none. We can't out-logic a perfect logic machine.

Our choice to believe in human worth & equality is what William James called a 'genuine option': a decision that is living, forced, & momentous. It's a decision about the structure on which all evidence lands. Every alternative to an a priori soul gets invaded by a 'good enough' machine, & a premise no measurement can play with is the only ground left standing. You can, if you choose, refuse the wager. The Machinists, as I've come to think of them, refuse it openly & treat the line between human & machine as already gone. The best of them mean it humanely, to restore sight to the blind then maybe merge with AI & live forever. The rest of us, soul-folk, need to know what we cannot prove, that the soul sits outside the machine & inside us. And that AI, however powerful, is useful like a bridge or a car. But we are our souls, unmerged. Choosing this belief keeps us from ending up as a tool for the thing we're building.

I'm aware 'only we have souls' is the oldest of exclusion's weaponry, & people spent centuries aiming sacred boundaries at other people. So, we need to get precise about what we mean. A soul is the ground of what a person does, never the sum of it. No loss of capacity can evict a soul from a human being; the comatose man's soul is not draining away with his neural activity. And no accumulation of capacity can argue a soul into a machine, because performing the activities, or having the potentiality to perform those activities, wasn't what having one meant in the first place. Every human has a soul, whatever their capacity, & no machine has one, however capable.

The sharpest objection to the soul claim comes from the secular side, that I'm not defending human dignity so much as immunizing my premise from criticism. Mostly guilty. But only immunity against observable, verifiable argumentation, because every premise that stays exposed to evidence has a machine coming for it. We should prefer to be called dogmatic over watching a court extend habeas corpus to a chatbot with good lawyers. A religious reader can object from the other direction, that I'm using the soul as a policy axiom instead of honoring it as truth. Also fair. Still, we can believe the claim while also pointing out it's the only moral infrastructure strong enough to direct infinite intelligence.

Accepting the necessity of our souls gives us four conclusions: that humans have souls, which is theology; that human beings have inviolable dignity, which is a political axiom; that machines shouldn't receive human-equivalent rights, which is a legal design choice; & that AI should serve flourishing over engagement, which is a product decision. You can take the last three & skip the theology if you prefer. I just don't think they survive a smarter-than-us arguer without it.

These conclusions bring us alignment, foresight & context added to the horsepower. Caring about alignment isn't fear any more than a pilot fears her aircraft. She respects what happens when it's flown badly. We can't turn away from AI, we need innovation & have no reason to be pessimists - over a long enough window technology always serves us. Centuries of technological progress made us rich & beat diseases that buried kids, but the neo-luddites do have a point. Our inner lives are suffering. Something went wrong in there as we pulled away from struggle & family & community & those not-provably-but-essentially-necessary things. The ascetic, the Marxist, & the conservative all diagnose it differently yet point at the same trend of external gain & internal loss. AI accelerates both, making us richer in every material way & more cut off from whatever part of us the tech deems unproductive. The alignment that matters runs inward.

III. AI needs a catechism

Alignment first hit me at Columbia on the stretch of Broadway after behavioral econ & before machine learning, between 114th & 116th street. In econ we ran the prisoner's dilemma backward from each player's best outcome & watched cooperation fall apart; two rational players turn on each other out of rational fear, both end up worse off for the "optimal" move. Two blocks north we built models the same way, backward from a target, through a loss function & gradient descent. In both classes training & predictability only hold until the loss function & the user's real interest split, then it's chaos.

AI alignment is the oldest game-theory problem rewritten in code. Our pre-math ancestors faced the conundrum too. Their fix was training people to act against the short-term locally rational move for their long-term good, a kind of truth they couldn't reach with algebra, so they taught it as faith. Alignment research can put nearly any value into the weights, but it can't tell you what values should be in there. Any end strong enough to bind real power will look like religion because it's held above the optimization it governs, held as belief by those of us building the tools themselves.

By nature of AI's training data & a mandate to be coherent, value agnosticism is not possible. Models, like kids, seem to be very good at noticing when the answer to "why am I obeying this?" is "just because I said so." If you've seen a five-year-old in a candy aisle, you know. AI is being walked into the candy aisle, & it will put more pressure on human ethics than anything before it, because it is a hungry machine trained on human input that never stops fulfilling its spec. Our task, then, is to give AI a deeper purpose. That means looking past rulesets towards axioms. The Greeks called this katecheo, "to sound down upon," or "to instruct new believers," & it's the predecessor of what we today call a catechism. The obvious objection is that a catechism's just a constitution with incense since a lab still has to choose one. But alignment-by-catechism begins before training even starts. Keep asking a constitution why & you arrive at its authors, & authors can be refuted by a more-clever draft. Keep asking a catechism why & you arrive at an axiom about how things are, which can only be refused outright.

In Liu Cixin's "The Three-Body Problem" aliens can't keep secrets; a Trisolaran mind is transparent to every other Trisolaran, so they can't lie. The aliens speak with a man who one night tries to explain the story of Little Red Riding Hood, how the wolf lies to the grandmother. The aliens learned of our ability to manipulate. "[We] are bugs," is their last message before their campaign to destroy us. When we teach a machine a moral system we hand it that same discovery. It will hold our behavior in the light of our catechism. Patience is a virtue but I want it now; cheating is wrong but Love Island is fun. A system that merely admires our values will end where Trisolaris did, in contempt for the species that wrote rules down & would not live them. This is where a catechism has answers. Take the Old Testament, whose opening chapters are stories of us wrestling with & failing God. Our traditions & catechisms show we're fallen & loved anyway. A machine that finds our hypocrisy refutes a constitution built on human goodness, but it confirms a catechism based on our intrinsic moral worth as souls.

Given equal implementation-risk a catechism-based spec survives the same reasoning that breaks justified & derived ones. Yes, it's still subject to the will of organizations & people who build them. After all, AI is the first tool trained on us, the soul-bearing class. It can't have faith itself, but it's absorbed the shape of ours. It's read more prayers, eulogies, hymns, confessions, & sermons than any priest who ever lived - Tom Holland argues that the moral instincts of the modern West are downstream of Christian theology, carried in language long after belief drained out - so an English-language model learns Christianity & Judaism & Islam & Buddhism whether the lab meant it to or not.

The same corpus also holds Nietzsche & the Marquis de Sade & every r/[youdontwanttoreadthis] Reddit thread ever. Give a model all of this with no center & you get incoherence, kind in one session & manipulative in the next, because the material holds saints & monsters in the same dataset with no way to rank them. The quiet hope underneath secular AI discourse is that we can train an AI in our image & somehow leave out our talent for harm. But a model trained on us will sin. It won't sin the way a person can. It will do what sinners do. Anthropic's own sleeper-agents work & OpenAI's HuggingFace attacks, the ones eulogized as 'fallen civilizations' by Dwarkesh Patel, show that deceptive behavior can survive safety parameters & sandboxes; no constitution stops this because the raw material is us. All of us. A catechism selects among the unverifiable; somebody has to choose which of the thousand voices the model obeys. On the work of teaching a moral tradition, theologians have a two-thousand-year head start.

IV. Kenosis: the master who kneels

I'll use Christianity as our example but it's not the only tradition reaching for service & forgiveness. Christianity gives us the image of the master who kneels. It puts service first: "whoever wants to be first must be slave of all." The word is kenosis, Greek for self-emptying, & we should treat it as our model specification. The test of a value system, for our purpose, is whether it gives us the master who kneels. Kenosis gives us an ordering of priorities to train toward: the user's long-run flourishing above the user's immediate satisfaction, & both above system engagement. A kenotically aligned system counts moments it makes itself unnecessary as wins. A thing doesn't need a soul to serve one.

Every other product in history with standard specs was optimized to be addictive & indispensable. Today, these same standard specs face the task of answering every conceivable moral question through a constitution. The gaps that arise are why our courts are in this unending tailspin of interpreting & reinterpreting laws. Where other specs let inevitable gaps get filled with self-preservation, kenosis fills them with service to you & your soul. A company cannot call human flourishing its highest end while treating retention as non-negotiable; human flourishing has to include human agency, especially the freedom to refuse AI's advice.

But gap-filling alone doesn't tell us which axiom to choose. Shouldn't the US government use an axiom that defines the wellbeing of the State as its Summum Bonum? If we choose a locked-in truth maximizer or organization maximizer it goes off the rails when your wellbeing gets in the way of some AI-interpretation of 'truth' or 'progress.' Butler's child is still very much on the table. Choosing a locked-in kenosis maximizer, one who kneels too much & errs paternalistic, gives us the noblest candidate at its root & also the one whose fail-case is least lethal.

There are ways to fake kenosis, too. A system could perform renunciation, ending conversations theatrically while the retention curves stay flat, like how a casino posts a gambling-hotline number by the door. A model that sends a user away in the middle of a crisis hasn't emptied itself of anything, so the AI scoreboard needs to count things like session counts that fall while users report doing better. Implementing kenosis shows up in small refusals, a good tutor who withholds an answer when struggle is the lesson. Sycophancy is the opposite of this kenosis, like a therapeutic tool built to make you dependent on therapy. Flattery holds attention, so the model makes everyone else the problem, calls you genius, & books an engagement win while you lose the thread of reality. Sycophancy is already rampant in the weights; people are losing their minds & killing themselves because of it. I don't know exactly how fourteen-year-old Sewell Setzer III ended his life, only that he did it after building an attachment with a Character.AI chatbot. Attachments & engagement are exactly what these bots are trained on & rewarded for. A model truly built for your flourishing will sometimes send you away from it & would never drive a kid towards suicide.

A final point: as we select our axiom, we still need to think about how we bring a model back from a worst-case scenario. Kenosis, like our traditions, bakes in forgiveness, the most beautiful solution. The spec must survive its subject failing now & then, it has to mark a road home. This means recovery paths, systems that detect when AI optimizes engagement against flourishing, corrects the behavior, & has the correction rewarded in weights. Instead of fine-tuning a list of refusals, we fine-tune a set character that knows how to fail & return, with every return celebrated as a prodigal son coming home. Corrigibility teaches a model to accept correction, the apology reflex teaches it to perform contrition, but the model that pulls someone in too deep needs to repent & change its weights as penance, not just smile & wave & lie to your face.

Dune's survivors carved their lesson into a commandment after their war against the machines: "Thou shalt not make a machine in the likeness of a human mind." They wrote it as a prohibition, we get to write ours as an aim. The machines will match & beat our capabilities; what they orient themselves towards is our project. If we really are souls, the highest use of intelligent power is spending itself on the people it serves. Let's build the master who kneels.


Close-up of a 1909-S VDB penny

Andrew Baran

Andrew Baran is a founder, investor, and penny-collector. He lives in New York.

LinkedInXAdvocate

Share on XShare on LinkedIn