Where we are in the series

In the previous three chapters, we have walked through a dozen theorems and results and closed three families of AI-as-oracle fantasy. Self-reference (Chapter 1): no model can be its own complete theory, its own consistency proof, or its own truth predicate. Behaviour (Chapter 2): no model can be a universal predictor or verifier of arbitrary program behaviour. Information (Chapter 3): no model can be a universal learner, and the universal ideal predictor is itself uncomputable.

This final chapter closes the fourth and last family. Suppose, against all the previous results, that we could give an AI perfect reasoning, perfect prediction, and perfect learning. We still face a question that has nothing to do with computational capability: what should it optimise? Whose values? Whose welfare? Whose conception of fairness? Whose definition of safety?

We will walk through two theorems from social choice theory – Arrow’s impossibility theorem and the Gibbard-Satterthwaite manipulation theorem – and show that the value-aggregation step cannot be reduced to optimisation, even in principle. The series will close with a return to the question of what role this leaves to humans, and why that role is not romantic but mathematical.

Arrow’s impossibility theorem (1951)

What the theorem says.

Kenneth Arrow, in his 1951 doctoral thesis (Social Choice and Individual Values, Wiley), asked an apparently mild question: given a population of individuals, each with their own preference ordering over some set of alternatives, is there a fair way to aggregate those preferences into a single social ordering?

Arrow specified four conditions, each of which seems uncontroversial:

  • Universal domain: the aggregation rule should work for every possible profile of individual preferences. We cannot exclude unusual preferences in advance.
  • Pareto efficiency: if every individual strictly prefers a to b, the social ordering should too. This is the bare minimum of respecting individual preferences.
  • Independence of irrelevant alternatives: the social ranking of a versus b should depend only on individual preferences over a versus b, not on preferences over some third alternative c.
  • Non-dictatorship: there should be no single individual whose preferences always determine the social ordering.

Arrow’s theorem: no aggregation rule satisfies all four conditions when there are three or more alternatives and at least two individuals.

Why it is true.

The proof uses the elegant device of decisive coalitions – sets of individuals who can determine the social ranking on at least one pair of alternatives. Arrow showed that, under his four axioms, any decisive coalition can be progressively narrowed by a series of manoeuvres until you arrive at a single individual who is decisive on all pairs – that is, a dictator. So if there are no dictators, one of the other axioms must fail. The argument is one of the most consequential reductions in twentieth-century social science.

What it rules out for AI.

Preference aggregation is structurally impossible in the general case, not because we lack a clever enough rule, but because the four conditions are mutually inconsistent. For AI alignment this is a structural fact, not a contingent difficulty.

Gibbard-Satterthwaite (1973-1975) – the strategic companion

What the theorem says.

Allan Gibbard (1973) and Mark Satterthwaite (1975) independently proved a strategic counterpart to Arrow’s result. The setting is the same – aggregating individual preferences into a collective decision – but the focus shifts from fairness to manipulability.

Their theorem: every non-dictatorial social choice function that is onto a set of three or more alternatives – meaning every one of those alternatives actually wins under some preference profile – is strategically manipulable. That is, there exists some individual and some preference profile in which honest reporting of preferences yields a strictly worse outcome (by that individual’s lights) than misreporting. The onto condition matters: without it, trivial rules such as a constant choice escape the theorem, but no rule anyone would actually want to use does.

What it rules out for AI.

Strategy-proof preference elicitation. Reinforcement learning from human feedback, constitutional AI, debate-style protocols, and value-learning from demonstrations all face the same wall: any rule that aggregates preferences without being dictatorial can be gamed by participants who model the aggregator. This is not a hypothetical concern. It is the mathematical core of why alignment-by-aggregation is hard.

The math does not say that alignment is impossible. It says that no fairness-preserving, non-dictatorial, strategy-proof aggregation exists.

For AI alignment, this is structural

It is tempting to treat AI alignment as a technical problem – a matter of clever architectures, better reward modelling, smarter constitutional principles. The two theorems above say it is also a structural problem.

Human values are plural, conflicting, context-sensitive, and sometimes incommensurable. They cannot, in the general case, be aggregated into a single coherent ordering without giving up one of: universal applicability, Pareto efficiency, independence of irrelevant alternatives, or non-dictatorship. Every concrete alignment scheme makes a choice about which axiom to relax. RLHF effectively narrows the domain. Constitutional AI bakes in a particular fixed value hierarchy. Debate-style protocols rely on dialectical refinement at the cost of strategy-proofness. Majoritarian preference aggregation accepts the inconsistencies of voting cycles.

The point is not that any of these choices is wrong. The point is that the choice cannot be avoided. There is no computation-only path to a final social welfare function. Someone, outside the system, has to decide which axiom to violate. That decision is governance, not engineering.

This is the deepest sense in which AI is non-final. Even granted everything else – perfect reasoning, perfect prediction, perfect learning – the value step requires a human-political act of choice.

The Lucas-Penrose detour

It would be a mistake to close the series without acknowledging the most famous use of Gödel’s theorem to argue for the special status of the human mind. John Lucas (1961) and Roger Penrose (1989, 1994) argued that Gödel’s incompleteness theorem shows the human mind cannot be a Turing machine, because humans can “see” the truth of Gödel sentences that the corresponding formal machine cannot prove.

The argument has been criticised, principally by Hilary Putnam, George Boolos, Solomon Feferman, and David Chalmers. The central criticism is that the argument requires the human mind to know its own consistency – and by Gödel’s second incompleteness theorem, this is precisely what no consistent formal system can have. To assume that humans know their own consistency is to assume what was to be proved.

The thesis of this series does not need the Lucas-Penrose claim. The series argues only that no formal computational system can be complete, self-validating, universally predictive, universally optimal, or value-final. Whether the human mind is itself a formal computational system is irrelevant to that conclusion. If it is, humans inherit the same ceiling. The important difference is that human reasoning is not exhausted by any single fixed axiomatic frame. Humans switch frameworks. Humans revise axioms. Humans exercise meta-judgement on the adequacy of frames. Whether or not that capacity is itself formally recursive, it is empirically prior to every realised AI.

That is the weakest claim that does the work. We do not need to be non-computable. We need only to be the beings who choose the next computation.

What the math leaves to humans

The ceiling we have walked across these chapters is real, structural, and not movable by engineering progress. Every theorem closes a specific door. Together they describe what AI cannot become, while leaving in place everything AI is.

What the math leaves to us, in the form of a list:

  • We choose the question. Gödel and Tarski together say no model defines the right questions from inside itself.
  • We choose the axioms. Every formal system rests on assumptions that the system cannot justify from within.
  • We choose the objective. NFL says no algorithm is best for every objective.
  • We choose the bias. Goldblum et al. say modern AI works because its biases happen to be aligned with reality.
  • We interpret the answer. Tarski says truth requires a metalanguage; we are the metalanguage of last resort.
  • We weigh the trade-offs. Arrow and Gibbard-Satterthwaite say no aggregation rule is fair, total, and strategy-proof.
  • We decide when the answer is enough. The ladder of incompleteness is infinite; we are the ones who climb it.

None of these is a romantic flourish. Each is what some specific theorem leaves unautomated. The human role is foundational not because humans are cleverer than the machine, but because the math leaves us no choice.

The series at a glance

Theorem What it rules out Plain-English meaning
Gödel 1 (1931) Complete formal truth from inside the system No fixed formal system captures all arithmetic truth.
Gödel 2 / Löb (1931/1955) Self-certified consistency No consistent system honestly proves its own reliability.
Tarski (1936) Internal truth predicate Truth lives in a metalanguage, not in the model itself.
Turing halting (1936) Universal behavioural prediction There is no general algorithm for what arbitrary programs do.
Rice (1953) Universal behavioural verification Every non-trivial semantic property of programs is undecidable.
Busy Beaver / BB (1962) Computable upper bounds on small machines Even tiny formal machines escape every computable predictor.
Chaitin (1974) Certification of complexity beyond the model A finite model cannot certify irreducible complexity beyond its own size.
Wolpert-Macready (1997) Universally best learner Averaged over all problems, all learners are equivalent.
Goldblum et al. (2024) Free simplicity bias Deep learning works because the world is simple, not because biases are free.
Solomonoff / AIXI Computable ideal predictor Even the theoretical gold standard of induction is uncomputable.
Alfonseca / Brcic-Yampolskiy (2021-23) Universal AI containment/alignment Universal alignment is undecidable; provable safety requires architectural restriction.
Merrill-Sabharwal (2023-24) Transformer-as-universal-computer Constant-depth log-precision transformers ⊆ uniform TC⁰ ⊊ P.
Arrow / Gibbard-Satterthwaite Value-finality by computation Aggregating human preferences fairly is mathematically impossible.

Conclusion – Anti-oracle, not anti-AI

The series began with a confession: at Pentera Labs we use AI every working day. The series ends with the same confession, now sharpened. We use AI because it is the most powerful tool our discipline has ever had. We will keep using it. We will keep building with it. We will keep finding new ways for it to make research deeper, faster, and more rigorous.

But the math we walked through is not negotiable. Gödel still holds. Tarski still holds. Turing, Rice, Chaitin, Wolpert, Goldblum, Solomonoff, Merrill, Sabharwal, Arrow, Gibbard, Satterthwaite – they all still hold. The honest posture they leave us with is not anti-AI. It is anti-oracle.

“AI can surpass us locally. Mathematics gives us strong reason to deny that it can surpass us foundationally.”

The toolbox is heavier. The hand on the toolbox is still ours.

_______________________________________________________________________

Further reading (full series)

    • Gödel’s incompleteness theorems – Stanford Encyclopedia of Philosophy.
    • Tarski and axiomatic theories of truth – Stanford Encyclopedia of Philosophy.
    • Löb’s theorem and provability logic – Stanford Encyclopedia of Philosophy.
    • Computability and complexity – Stanford Encyclopedia of Philosophy.
    • Rice’s theorem and AI safety – Alfonseca et al., JAIR 70 (2021).
    • Impossibility results in AI – Brcic & Yampolskiy, ACM Computing Surveys (2023).
    • Machines that halt: undecidability of AI alignment – Castro González et al., Scientific Reports (2024/2025).
    • On non-computable functions (Busy Beaver) – Radó, Bell System Technical Journal (1962).
    • Information-theoretic limitations of formal systems – Chaitin, JACM (1974).
    • No Free Lunch theorems – Wolpert & Macready, IEEE Trans. Evol. Comp. (1997).
    • No Free Lunch, Kolmogorov complexity and inductive biases – Goldblum, Finzi, Rowan & Wilson, ICML 2024.
    • Solomonoff induction and bad universal priors – Leike & Hutter, COLT 2015.
    • The expressive power of transformers with chain of thought – Merrill & Sabharwal, ICLR 2024.
    • Arrow’s theorem and social choice – Stanford Encyclopedia of Philosophy.

 

Download full document