Skip to main content
ExplainerMachine EthicsExplainer· 3 min read· in Culture

How the 'Veil of Ignorance' is Shaping the Rules of Artificial Intelligence

A 1971 philosophical thought experiment by John Rawls has become a foundational blueprint for programming fairness into modern artificial intelligence systems.

By Joao Marques

Procedural Ethicists 40%Utilitarian Engineers 35%Algorithmic Skeptics 25%
Procedural Ethicists
Argue that fairness in AI can only be achieved by blinding the system to demographic variables during training.
Utilitarian Engineers
Focus on maximizing overall positive outcomes and minimizing harm, often relying on crowdsourced preference data.
Algorithmic Skeptics
Warn that mathematical approximations of justice still reflect the biases of the developers who write the rules.

Perspectives this story doesn't cover

  • Marginalized communities whose data is excluded from the training sets
  • Legal scholars debating liability for autonomous decisions

At a glance

  1. Artificial intelligence developers are increasingly using John Rawls's 'veil of ignorance' to design fair behavioral constraints for autonomous systems.
  2. The concept requires decision-makers to establish rules without knowing their own demographic identity, naturally incentivizing the protection of vulnerable groups.
  3. Techniques like Constitutional AI apply this by using a blinded 'AI judge' to enforce high-level principles during the reinforcement learning phase.
  4. Critics warn that while the procedure mimics philosophical fairness, the underlying rules are still written by a small group of corporate engineers.

Why it matters now

As artificial intelligence systems increasingly make decisions about healthcare, hiring, and navigation, the ethical frameworks hardcoded into their models will determine who benefits and who gets left behind. Understanding the philosophy behind these systems demystifies how machines are being taught to navigate human morality.

When engineers at major artificial intelligence labs sit down to write the behavioral constraints for their models, they are no longer just writing code. They are deciding whose values win when a machine has to make a choice, and they are doing so with global deployment on the line. To figure out what those rules should be, developers are increasingly reaching for a thought experiment published by political philosopher John Rawls in 1971.[1][2]

Rawls introduced the concept of the "original position," a hypothetical scenario designed to strip away human bias. He asked readers to imagine designing a society's rules from behind a "veil of ignorance"—a state where no one knows their own race, wealth, gender, or natural endowments. If you do not know whether you will be born rich or poor, you are highly likely to design a system that protects the most vulnerable. As the Stanford Encyclopedia of Philosophy notes, this "occludes information, for instance, about principals' age, sex, religious beliefs," rendering the problem of choice determinate.[1]

How theoretical philosophy translates into a two-phase machine learning pipeline.

Historically, this remained a purely academic exercise. But the rise of autonomous systems has turned the veil of ignorance into a practical engineering requirement. When the Massachusetts Institute of Technology launched the "Moral Machine" experiment in January 2016, it collected human decisions on how self-driving cars should allocate risk during unavoidable crashes, testing respondents across more than 200 countries. The data revealed deep cultural biases, proving that crowdsourced morality is highly subjective.[3]

To prevent an artificial intelligence from simply memorizing and amplifying these demographic biases, researchers have pioneered a technique called "Constitutional AI" (CAI). As synthesized by the Factlen Editorial Team, the process operates in exactly two phases: Supervised Learning (SL-CAI) and Reinforcement Learning (RL-CAI). Instead of having human labelers grade every possible response, the developers write a constitution of high-level principles. The AI then grades its own outputs against that constitution.[5]

Researchers are attempting to quantify moral stability using a 0-to-1 scoring system.
As synthesized by the Factlen Editorial Team, the process operates in exactly two phases: Supervised Learning (SL-CAI) and Reinforcement Learning (RL-CAI).

Here, the veil of ignorance moves from theory to architecture. During the reinforcement learning phase, the AI judge evaluating the model's behavior operates completely blind to the identity of the user. It only knows the principles in its constitution. Models are evaluated on a Stability Score between 0 and 1 that aggregates exactly three core dimensions of welfare: productivity, survival, and conflict.[5]

The stakes for getting this right are immense. By September 2025, researchers at Harvard University were testing "Inverse Constitutional AI," a method that attempts to reverse-engineer a fair constitution from diverse human preferences using approval voting algorithms. The goal is to mathematically compress human values into a set of rules that a machine can follow, without letting the majority tyrannize the minority.[5]

The physical infrastructure where procedural fairness is encoded into neural networks.

Yet, this approach carries profound limitations. A machine operating behind a veil of ignorance is still executing a constitution written by a small group of developers. While some labs have experimented with crowdsourcing these rules through a project called "Collective Constitutional AI," critics point out that democratic input does not automatically resolve moral conflicts. Participatory approaches often overlook questions of enforceability, leaving the final arbitration to the engineers.[5]

The translation of political philosophy into neural networks represents a fundamental shift in how society regulates technology. We are no longer relying on Isaac Asimov's 1942 "Three Laws of Robotics," which imagined natural language rules for physical robots. Instead, we are encoding the procedural fairness of the original position directly into the weights of the model. The next time a language model refuses a prompt, the decision will be based on a mathematical approximation of justice.[2][5]

Terms to know

Original Position
A hypothetical scenario in political philosophy where individuals agree on the rules of society without knowing their own place within it.
Constitutional AI
An alignment technique that uses a written set of principles to guide an AI model's behavior during its training phase.
Reinforcement Learning
A machine learning process where a model learns to make decisions by receiving rewards or penalties based on its actions.
Trolley Problem
A classic ethical dilemma asking whether it is morally permissible to sacrifice one person to save a larger number of people.

Questions readers ask

What is the veil of ignorance?

It is a philosophical thought experiment created by John Rawls in 1971. It asks people to design the rules of society without knowing their own race, gender, or wealth, theorizing that this blindness leads to fairer rules.

What is Constitutional AI?

It is a training method where an AI model is given a 'constitution' of high-level principles. Instead of humans grading the AI's responses, a secondary AI judge evaluates the outputs against these principles to ensure safe behavior.

How does the Moral Machine relate to this?

The Moral Machine was an MIT experiment that crowdsourced human decisions on the Trolley Problem for self-driving cars. It highlighted the difficulty of programming ethics, as different cultures strongly disagreed on who should be saved in a crash.

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Procedural Ethicists 40%Utilitarian Engineers 35%Algorithmic Skeptics 25%
  1. [1]Stanford Encyclopedia of PhilosophyProcedural Ethicists

    Original Position

    Read on Stanford Encyclopedia of Philosophy
  2. [2]Stanford Encyclopedia of PhilosophyProcedural Ethicists

    John Rawls

    Read on Stanford Encyclopedia of Philosophy
  3. [3]WikipediaUtilitarian Engineers

    Moral Machine

    Read on Wikipedia
  4. [4]WikipediaUtilitarian Engineers

    Trolley problem

    Read on Wikipedia
  5. [5]Factlen Editorial TeamAlgorithmic Skeptics

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Culture stories with full source coverage and perspective breakdowns delivered to your inbox.