How the 'Veil of Ignorance' is Shaping the Rules of Artificial Intelligence
A 1971 philosophical thought experiment by John Rawls has become a foundational blueprint for programming fairness into modern artificial intelligence systems.
By Joao Marques
- Procedural Ethicists
- Argue that fairness in AI can only be achieved by blinding the system to demographic variables during training.
- Utilitarian Engineers
- Focus on maximizing overall positive outcomes and minimizing harm, often relying on crowdsourced preference data.
- Algorithmic Skeptics
- Warn that mathematical approximations of justice still reflect the biases of the developers who write the rules.
Perspectives this story doesn't cover
- Marginalized communities whose data is excluded from the training sets
- Legal scholars debating liability for autonomous decisions
At a glance
- Artificial intelligence developers are increasingly using John Rawls's 'veil of ignorance' to design fair behavioral constraints for autonomous systems.
- The concept requires decision-makers to establish rules without knowing their own demographic identity, naturally incentivizing the protection of vulnerable groups.
- Techniques like Constitutional AI apply this by using a blinded 'AI judge' to enforce high-level principles during the reinforcement learning phase.
- Critics warn that while the procedure mimics philosophical fairness, the underlying rules are still written by a small group of corporate engineers.
Why it matters now
As artificial intelligence systems increasingly make decisions about healthcare, hiring, and navigation, the ethical frameworks hardcoded into their models will determine who benefits and who gets left behind. Understanding the philosophy behind these systems demystifies how machines are being taught to navigate human morality.
When engineers at major artificial intelligence labs sit down to write the behavioral constraints for their models, they are no longer just writing code. They are deciding whose values win when a machine has to make a choice, and they are doing so with global deployment on the line. To figure out what those rules should be, developers are increasingly reaching for a thought experiment published by political philosopher John Rawls in 1971.[1][2]
Rawls introduced the concept of the "original position," a hypothetical scenario designed to strip away human bias. He asked readers to imagine designing a society's rules from behind a "veil of ignorance"—a state where no one knows their own race, wealth, gender, or natural endowments. If you do not know whether you will be born rich or poor, you are highly likely to design a system that protects the most vulnerable. As the Stanford Encyclopedia of Philosophy notes, this "occludes information, for instance, about principals' age, sex, religious beliefs," rendering the problem of choice determinate.[1]
Historically, this remained a purely academic exercise. But the rise of autonomous systems has turned the veil of ignorance into a practical engineering requirement. When the Massachusetts Institute of Technology launched the "Moral Machine" experiment in January 2016, it collected human decisions on how self-driving cars should allocate risk during unavoidable crashes, testing respondents across more than 200 countries. The data revealed deep cultural biases, proving that crowdsourced morality is highly subjective.[3]
To prevent an artificial intelligence from simply memorizing and amplifying these demographic biases, researchers have pioneered a technique called "Constitutional AI" (CAI). As synthesized by the Factlen Editorial Team, the process operates in exactly two phases: Supervised Learning (SL-CAI) and Reinforcement Learning (RL-CAI). Instead of having human labelers grade every possible response, the developers write a constitution of high-level principles. The AI then grades its own outputs against that constitution.[5]
As synthesized by the Factlen Editorial Team, the process operates in exactly two phases: Supervised Learning (SL-CAI) and Reinforcement Learning (RL-CAI).
Here, the veil of ignorance moves from theory to architecture. During the reinforcement learning phase, the AI judge evaluating the model's behavior operates completely blind to the identity of the user. It only knows the principles in its constitution. Models are evaluated on a Stability Score between 0 and 1 that aggregates exactly three core dimensions of welfare: productivity, survival, and conflict.[5]
The stakes for getting this right are immense. By September 2025, researchers at Harvard University were testing "Inverse Constitutional AI," a method that attempts to reverse-engineer a fair constitution from diverse human preferences using approval voting algorithms. The goal is to mathematically compress human values into a set of rules that a machine can follow, without letting the majority tyrannize the minority.[5]
Yet, this approach carries profound limitations. A machine operating behind a veil of ignorance is still executing a constitution written by a small group of developers. While some labs have experimented with crowdsourcing these rules through a project called "Collective Constitutional AI," critics point out that democratic input does not automatically resolve moral conflicts. Participatory approaches often overlook questions of enforceability, leaving the final arbitration to the engineers.[5]
The translation of political philosophy into neural networks represents a fundamental shift in how society regulates technology. We are no longer relying on Isaac Asimov's 1942 "Three Laws of Robotics," which imagined natural language rules for physical robots. Instead, we are encoding the procedural fairness of the original position directly into the weights of the model. The next time a language model refuses a prompt, the decision will be based on a mathematical approximation of justice.[2][5]
Terms to know
- Original Position
- A hypothetical scenario in political philosophy where individuals agree on the rules of society without knowing their own place within it.
- Constitutional AI
- An alignment technique that uses a written set of principles to guide an AI model's behavior during its training phase.
- Reinforcement Learning
- A machine learning process where a model learns to make decisions by receiving rewards or penalties based on its actions.
- Trolley Problem
- A classic ethical dilemma asking whether it is morally permissible to sacrifice one person to save a larger number of people.
Questions readers ask
What is the veil of ignorance?
It is a philosophical thought experiment created by John Rawls in 1971. It asks people to design the rules of society without knowing their own race, gender, or wealth, theorizing that this blindness leads to fairer rules.
What is Constitutional AI?
It is a training method where an AI model is given a 'constitution' of high-level principles. Instead of humans grading the AI's responses, a secondary AI judge evaluates the outputs against these principles to ensure safe behavior.
How does the Moral Machine relate to this?
The Moral Machine was an MIT experiment that crowdsourced human decisions on the Trolley Problem for self-driving cars. It highlighted the difficulty of programming ethics, as different cultures strongly disagreed on who should be saved in a crash.
Sources
[1]Stanford Encyclopedia of PhilosophyProcedural EthicistsOriginal Position
Read on Stanford Encyclopedia of Philosophy →
[2]Stanford Encyclopedia of PhilosophyProcedural EthicistsJohn Rawls
Read on Stanford Encyclopedia of Philosophy →
[3]WikipediaUtilitarian EngineersMoral Machine
Read on Wikipedia →
[4]WikipediaUtilitarian EngineersTrolley problem
Read on Wikipedia →
[5]Factlen Editorial TeamAlgorithmic SkepticsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Culture
See all →Orthography
Decoding the Page: How the Four Major Writing Systems Translate Sound into Sight
7 sources
Cultural Transmission
How Ideas Cross Borders: The Four Pathways of Cultural Diffusion
6 sources
Structural Engineering
Gravity and Geometry: How Hoop Stress and Thrust Lines Keep Historical Domes Standing
9 sources
Wildlife Photography
Audubon Photography Awards Grand Prize Goes to 'Ghost of the Desert' Wildlife Shot
6 sources
Every angle. Every day.
Get Culture stories with full source coverage and perspective breakdowns delivered to your inbox.




