Menu

Research Epistemology 01

Epistemology

Intelligent Epistemology

Consistent inference exists. The rest is what the grounds carry.

Three horns, one existence proof · Prop. 19.1

Field
Epistemology
Document
MU and Epistemic Zero
Length
81 pages
Edition
Interactive, 2026

The paper (PDF) The book

Abstract

MU is the paper's name for one principle: consistent inference is possible. It is the thinnest fact reasoning can stand on. One line of reflexivity proves it. Everything above that floor turns on the word from: a conclusion follows from its grounds exactly insofar as every answer-relevant dependence is carried by those grounds. Labels, coordinates, search order and priors earn inferential force in one way only. The problem names them. Assume nothing the constraints omit. Surrender nothing they deliver.

Interactive research edition · six parts, thirty-three sections · every claim carries its § or theorem number

The argument

Thirty-three sections, folded into sixteen.

Nothing is dropped. Each row names the paper part, the sections this edition makes of it, and how many numbered statements and argument blocks sit underneath. From a one-line existence proof to the norms of inquiry.

  1. I MU and the Logic of Inference §1–§3 · 18 statements Consistent inference exists. One line of reflexivity proves it. Dependence on grounds then unfolds what any inference commits itself to, and Epistemic Zero states the rule in both directions.

    Everything turns on one word The answer depends on the problem and on nothing else

  2. II The Ground and Its Limits §4–§6 · 4 statements Logic, content, and an operator: an inference episode needs all three. MU is proved before it is applied. Gödel, Turing, Tarski and Arrow already supply the shapes a complete result can take.

    A proof is not an event of proving An empty class is a finished answer

  3. III The Mathematics of Rational Belief §7–§18 · 89 statements Four laws, each unique in its domain. Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion and revision are the same divergence read from two references. Values and costs then govern choice, and channels carry belief’s connection to the world.

    Three cuts, and then four laws Probability is not chosen. It is forced. Count every change exactly once One divergence, two reference points Sometimes the answer is a set, and that is the answer The kernel is where the world gets in

  4. IV Classical Problems: Resolutions, Reconstructions, and Boundaries §19–§23 · 31 statements · 31 argument blocks Regress, the criterion, Carroll’s tortoise, indifference, Goodman, induction, Gettier cases, disagreement and scepticism each become one question: what full answer do their grounds support? Some resolve. Some return a family. Some come back with a certified boundary.

    The trilemma has a fourth case A right answer can still be bad inference Thirty-one problems, and what follows from each

  5. V Science, Computation, and Society §24–§29 · 10 statements · 15 argument blocks Falsification, underdetermination, and the quotient a theory change preserves. A bounded reasoner pays for computation, and that price enters the problem rather than excusing the answer. Testimony, causal models and machine inference each run on a channel the problem has to name.

    Where inquiry gets done

  6. VI Norms and the Value of Knowledge §30–§33 · 5 statements · 3 argument blocks What an inference owes, and what knowing is worth. Internal support and external success are separate coordinates: a demon-world subject can be impeccable on the first and fail the second.

    The standards arrive with the claim

Part I

MU and the Logic of Inference

Part I

When you know a thing, to hold that you know it; and when you do not know a thing, to allow that you do not know it: this is knowledge.

Confucius, Analects 2.17, tr. Legge

Before the argument

One fact at the base, and everything it turns out to contain.

MU begins with the thinnest fact on which reasoning can stand: consistent inference is possible. The proof takes one line, since for any consistent proposition P, reflexivity gives P ⊢ P. Everything else turns on the relation the word from names: a conclusion follows from its grounds exactly insofar as every answer-relevant dependence is carried by those grounds.

Intelligent Epistemology, abstract

MU ≡ ∃I CI(I)

F = F̂ ∘ q

MU : Π ⟼ Answer(Π)

The paper boxes all three. The first is the floor, the second is what any inference commits itself to, the third is the operator the first two produce.

The paper runs 33 numbered sections in 6 parts over 81 pages, closing on an unnumbered conclusion. This edition maps them into 16 sections in six, and every section prints the paper range it covers in its eyebrow. Underneath sit 157 numbered statements and remarks, together with 49 argument blocks. Not one argument block appears before §19.

How claims are graded on this site

Results are graded by their statement form throughout: theorems, propositions, and corollaries are proved; arguments are defended on stated grounds without claim of proof; remarks locate, delimit, and connect.

Intelligent Epistemology, §1.1
  • proved Theorem, proposition, corollary, or lemma. Proved.
  • argued An argument block. Defended on stated grounds, without claim of proof.
  • located A remark. It locates, delimits, and connects.
  • defined A definition. It fixes a term and asserts nothing.
  • open The paper marks this one unsettled.

That sentence is the whole of the grading, and this site copies it rather than inventing a scale. A claim carries the register of the statement form it came from and no other. An argument summarised here stays argued however persuasive the argument is. Every tagged figure names the section, theorem, proposition, corollary, lemma, definition, or remark it comes from, and every number links to the page that states it.

What MU returns, object by object

consistent inference Def 1.3 The floor. At least one valid passage from a consistent premise set exists, and the identity instance exhibits it.
the inference problem Def 3.1 Seven slots: representation language, candidate domain, premises and constraints, consequence relation, question, required form of answer, and the rule for when two outputs count as the same.
the full answer Def 3.3 Everything the problem supports, together with the equivalences it declares. A point, an orbit, a family, or a certified empty class.
the inference episode Thm 4.1 Rule, content, and operator. Remove any one and no episode remains, though the abstract relation survives without the third.
the certificate Prop 6.1 An empty admissible class is a complete classification when the proof of emptiness comes with it. Gödel, Turing, Tarski, Fitch, and Arrow all return one.
probability Thm 9.2 MU's unique form for determinate scalar plausibility on a Boolean event structure, up to order-preserving regraduation.
relative entropy Thm 10.5 The unique measure of total informational change whose staged and flat descriptions balance, up to positive scale.
the two projections Remark 10.2 One divergence, two reference objects. Completion projects from a permitted reference; revision projects from the current state.
the credal family Def 11.4 The set of point-valued answers a problem permits. Where it holds several members, the set itself is the answer.
the evidence channel Def 16.1 The stochastic kernel that carries the connection to the world. A single reliability number is only a special summary of it.

This paper prints no symbol table. It introduces its vocabulary where it needs it, in definitions, and the list above is that vocabulary in the order the argument builds it. The rail beside you tracks the same ten objects, lighting each as the section that declares it arrives.

These minutes are counted off the page rather than guessed. The skim layer — headline, dek, plain box and verdict, sixteen times — runs about 10 minutes. The argument itself runs about 60. Opening the 6 technical blocks and the 8 proof ideas adds roughly 10 more. There is no layer control on the page. These are depths of reading rather than modes to select, and the paper prints no reading guide to build one from. The contents plate above is the shortest route in: six part rows, each naming the paper range it folds.

1 line of proof
For any consistent proposition P, reflexivity gives P ⊢ P. The premise set is consistent, the instance is valid, and existential introduction discharges the chosen P. Everything in the paper stands above that one line.
4 laws, none with an alternative
Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Revision is KL projection. Within their stated domains, no consistent alternative exists.

p. 3 · §1–§2 — the paper (PDF), opens in a new tab · The word from · Part I 01 / 16

Register — the paper's own margin letters:

declares ∃I · consistent inference p. 3 · Def 1.3 — the paper (PDF), opens in a new tab

The argument starts here The answer depends on the problem and on nothing else

Everything turns on one word

Every argument carries a promise: the conclusion comes from its grounds. The paper takes that promise at face value. If an answer is presented as following from stated grounds, then any feature that changes the answer must appear among those grounds or stay visible as unresolved freedom. Beneath the demand sits a floor thin enough to prove in a line. Consistent inference is possible.

In plain words When you say a conclusion follows from your reasons, you are making a claim you can be held to. Anything that changes your answer has to be one of the reasons you gave. If it is not, you were carrying something you never declared. The whole book is that one demand, taken seriously.

Watch the derivation assemble · film 1

Introduction, The Question in the Rain — The First Principle Chapter Three, The Gallery — The First Principle Chapter Six, Self-Grounding — The First Principle Chapter Thirteen, The Sceptic’s Self-Defeat — The First Principle Coda, The Return — The First Principle

The turn

A theory of principle

There are two ways to write a theory of reasoning. A constructive account builds inference out of proposed mechanisms and asks whether the machine works. A theory of principle starts from a constraint every genuine inference must satisfy and asks which forms meet it. This paper takes the second road. The road has a shape: state the problem, determine whether an answer exists, classify every answer up to the relevant equivalence, and require that a mere redescription leave the result unchanged.

The paper grades itself as it goes. Theorems, propositions, and corollaries are proved. Arguments are defended on stated grounds with no claim of proof. Remarks locate, delimit, and connect. That sentence in §1.1 is where this site gets its register marks. No claim here is graded higher than the statement form it came from.

What the structure says

The word from

Consider what the ordinary word is doing. An answer presented as following from its grounds has made a commitment about dependence. The commitment is checkable. Take any feature that would change the answer. It appears among the grounds, or the answer had help.

Labels, coordinates, orderings, priors, and preferences are the usual stowaways. Each is harmless once written down and given a role. None is harmless while it sits outside the stated problem and steers the result anyway. The commitment attaches to the relation rather than to any particular vocabulary for it, which is why the argument survives translation into other formal settings.

What the structure says

What MU says, and what it does not

MU is the principle that consistent inference is possible. In symbols it says only that there exists an I such that I is a valid inference from a consistent premise set. It does not say what follows from what. It does not privilege a logic, a language, or a method. It claims existence. Nothing more.

That thinness is deliberate. A floor thick enough to be interesting would be thick enough to be denied. The paper wants a base nothing built above it can undermine. Definition 1.1 keeps the relation apart from the episode: the formal core concerns the inferential relation, while claims about agents concern actual uses of it. The distinction pays off two sections later.

What the structure says

The proof, and the shape of the denial

Let P be any consistent proposition. Reflexivity gives the valid instance P ⊢ P. Its premise set is a single consistent proposition, so the instance is a consistent inference, and existential introduction discharges the chosen P from the conclusion. The theorem is about the kind rather than the witness. That is the whole proof.

The corollary runs the same machinery backwards. Suppose a language contains the proposition that no consistent inference exists. Suppose further that the proposition were consistent. Reflexivity would immediately make the identity inference from it to itself a consistent inference. The denial exhibits what it denies. This is stability under self-reference, not a rhetorical trap: the proof of the theorem never needed the denial, and the corollary only records what happens when someone tries.

Two movements of one principle follow: existence, then form. The MU Theorem establishes that a consistent inference exists. The Dependence Theorem, three pages later, unfolds what any such passage commits itself to once its conclusion is said to come from its grounds. Everything quantitative in the paper is downstream of that second movement. Existence first. Form after.

1 line of proof
For any consistent proposition P, reflexivity gives P ⊢ P. The premise set is consistent, the instance is valid, and existential introduction discharges the chosen P. That is the whole proof of the floor.
0 consistent denials
Suppose ¬MU were consistent. Reflexivity would then make the identity inference from ¬MU to ¬MU a consistent inference, witnessing the existence the proposition denies.
2 hypotheses
A reflexive inferential relation, and a language containing at least one consistent proposition. The theorem asks for nothing else.
MU is the floor: consistent inference exists.
Intelligent Epistemology, §1.2

Verdict The floor is not a postulate. Theorem 2.1 exhibits a consistent inference rather than assuming one, and Corollary 2.2 shows the denial cannot be consistently asserted, because reflexivity turns the denial itself into the witness it denies. Everything after this is what the demand contains once a domain names its objects.

From the paper
An answer that changes with an unstated label, coordinate, ordering, prior, or preference draws on more than the grounds supplied.

§1.2 · The word from

The whole discipline in one sentence, stated before any machinery arrives. Every later theorem is this demand met inside a domain that has named its objects.

The proof idea · Thm 2.1argued

Existence of consistent inference

HypothesesA reflexive inferential relation whose language contains at least one consistent proposition.

  1. Let P be a consistent proposition. The hypothesis says there is one.
  2. Reflexivity gives the valid instance P ⊢ P.
  3. Its premise set is {P}, which is consistent by choice of P.
  4. So P ⊢ P is a consistent inference, and existential introduction discharges P.
  5. The result is about the kind, not the witness: the chosen P disappears from the conclusion.

p. 4 · Thm 2.1 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

The companion volume
Assume nothing beyond what the constraints demand.

The First Principle · Reasoning (in) the Intelligence Age

The trade book runs the same argument for a reader with no formal training, opening on the same Confucius line this paper's Part I carries. It grades its claims in three bands rather than five: proved, best explanation, avowed.

Read the book

proved The proof, at full length One consistent proposition, one application of reflexivity, one existential introduction. Then the denial, which reflexivity turns into the witness it denies. ≈ 20s

Source: Intelligent Epistemology · Thm 2.1 · Cor 2.2 · Prop 19.1 · film 1

Plate 1

The proof at full length, and the denial that supplies its own witness. One consistent proposition, one application of reflexivity, one existential introduction (Thm 2.1, Cor 2.2). The fourth panel walks the Münchhausen horns and shows why none of them is this (Prop 19.1).

What the proof commits you to · The answer depends on the problem and on nothing else

p. 4 · §3 — the paper (PDF), opens in a new tab · The method · Part I 02 / 16

Register — the paper's own margin letters: ·

declares Π · the inference problem p. 4 · Def 3.1 — the paper (PDF), opens in a new tab

Everything turns on one word A proof is not an event of proving

The answer depends on the problem and on nothing else

Write the working presentation of a problem and the problem itself as two different things, joined by a map that forgets labels, coordinates, enumeration order, and temporary gauges. An answer rule depends only on the problem exactly when it factors through that map. Theorem 3.1 turns an ordinary word into an equation, and the equation generates the rest of the paper.

In plain words Two people can write the same problem down in different ways. If their answers differ, something in how they wrote it is doing work. That something was never declared. The fix is to say what counts as the same problem, then require the answer to depend only on that. Once you do, most of the argument writes itself.

Chapter Four, Epistemic Zero — The First Principle

Four shapes of a complete result

A complete answer is whatever the problem supports, kept whole. There are four shapes it can take. This card opens on the fourth: an optimum the case approaches and never reaches. Move the slider and the divergence falls toward an infimum no member attains. The other chips declare a symmetry, supply a second reference, or tighten a constraint until the answer changes shape.

the open constraint · shape (iv) · D = 0.130812 nats at p₁ = 0.7500 · infimum 0, attained by nobody
a member of the open case · p, with p₁ > ½ 0.750 1 0.250 2 excluded boundary
Definition 3.4, four shapes. This problem returns:
  • (i) one class, one representative, fixed by every symmetry
  • (ii) one orbit, its representatives exchanged by the symmetry
  • (iii) more than one inequivalent answer, retained as the family
  • (iv) an empty class or unattained optimum, with its certificate
  • members inequivalent answers retained
  • H(p) 0.5623 nats, at the position shown
  • DKL(p‖μ) 0.1308 departure from the uniform reference
  • boundary (½, ½) the point no member of the class reaches

The infimum of D(p‖μ) over the open case is 0, approached as p₁ falls to ½ and attained by no member, because the minimising point lies on the excluded boundary. At p₁ = 0.7500 the divergence reads 0.130812 nats, computed from the closed form and again from the sum over both atoms. The two disagree by exactly zero. Push the slider down and the number falls without ever arriving. The complete answer is three things: the case, the infimum 0, and the boundary point (½, ½) that no member reaches.

Shape is not a matter of how hard the problem is. It is a matter of what the problem carries. Every move that turns a family into a point on this card is a declaration. The card names it.

defined the four shapes are Definition 3.4 and the orbit rule is Corollary 3.3; the four presets are the paper's own worked examples in §11.4, recomputed here rather than quoted. argued the translation case is drawn on a finite window of at most 200 cells. That clamp is the frame's, not the paper's: on ℝ the invariant class is Lebesgue, the total mass is infinite at every finite window and beyond it, and no normalised invariant probability exists at all.

The measured residual
closed form against the atom sum
0

Each of these is one quantity computed two ways from the state on screen, with neither route derived from the other. A figure of 0 means the two routes landed on the same double. Anything else is the width of one. Where a line above reports how far two readings disagree, this is the figure it stands for.

Source: Intelligent Epistemology §3 · §11.4 · Def 3.4 · Cor 3.3 · Thm 11.3

Plate 2

Move the constraints and the symmetry, and watch the answer take one of Definition 3.4's four shapes. The four settings on the right are the paper's own worked examples from §11.4: a symmetric die, two standing references, translation on the line, and an open constraint whose optimum is never attained.

The turn

Writing the problem down

A problem is more than its data. It is a specification. It includes the language the grounds are expressed in, the candidates on offer, the constraints, the rule of consequence, the question being asked, the form of reply required, and the rule for when two outputs count as the same answer. Definition 3.1 makes all seven explicit and calls the whole specification an inference problem.

Any one of the seven may be stipulated, inherited from a domain, or supported by a further argument. Writing it down identifies its role and exposes whatever supports it. A language, consequence relation, reference measure, output form, or equivalence offered as a conclusion becomes the answer to a higher-order problem. It follows the same rule.

What the structure says

The equation

Let 𝒳 be the class of working presentations and 𝒫 the class of problems, joined by a surjection q that forgets everything the problem treats as irrelevant. An answer rule assigns to each presentation an answer in the answer space of its problem. Theorem 3.1 then says something short. The rule depends only on the problem exactly when it is constant on every class q identifies. That is exactly when it factors as F = F̂ ∘ q.

One direction is immediate. The other constructs F̂ by picking any presentation of a given problem and reading off its answer, then uses constancy to show the choice does not matter. A second condition rides alongside. Across a structure-preserving redescription, answers must transport with it. Mathematicians call that naturality.

Put two reasoners in front of the same problem. A difference in their answer spaces reveals a difference in grounds, representation, or procedure. Naming the difference puts it in the problem. Until it is named, the full answer keeps every possibility the shared grounds support. Nothing is discarded on trust.

What the structure says

What completeness costs

Definition 3.3 fixes the other half. A report is complete when it preserves the entire answer space, together with the equivalences the problem declares. Completeness is exhaustion relative to the question. Computability, decidability, and single-valuedness are separate properties. None is required.

Theorem 3.2 then closes the loop. A result presented as complete preserves every satisfying answer, every distinction the equivalence leaves inequivalent, and every unresolved family. A question, output form, equivalence, or selector compresses the result precisely when it makes the corresponding distinction irrelevant. Those data define the richer problem whose complete answer is the compressed one. Nothing is lost by insisting on this. Something is gained: the compression becomes visible as a further piece of structure rather than a silent convenience.

What the structure says

Canonical points and the symmetric pair

A point is canonical exactly when the problem contains structure that distinguishes and fixes it. Suppose a symmetry of the problem exchanges two satisfying answers. Every point rule that follows from the problem is invariant under that symmetry, so the invariant output retains the exchanged answers together. The full answer is the pair. Not one of them.

Enlarging the problem by an order supplies the distinction that selects one of them. The order is then part of the record. Its own standing depends on its own grounds. This is a small result with a long reach: it is why the paper never has to choose between a point and a family. The grounds choose.

The turn

Epistemic Zero, and the axiomatic method as a special case

The rule is bilateral. Two halves. Add only supported structure. Preserve every supported distinction. In plain words: assume nothing beyond what the constraints demand, and surrender nothing they deliver. Drop the first half and unstated assumptions creep in; drop the second and answers get compressed into points nobody paid for.

The ordinary axiomatic method falls out as an instance. Specify a language, an axiom theory, a consequence relation, a question, a form of answer, and an equivalence, and the complete theorem set is deductive closure. Nothing exotic happens here. The general rule becomes the familiar one because the familiar one already satisfies it.

The four shapes a complete result can take (Def. 3.4)
The four shapes a complete result can take (Def. 3.4)
ShapeWhat the problem has doneWhat the report must carry
A pointIts structure distinguishes and fixes one representative under every symmetrythe representative, and the structure that fixed it
An orbitA symmetry fixes the answer-class while leaving its representatives equivalentthe orbit, not one member of it
A familyMore than one inequivalent answer survives the stated groundsthe whole family, as the answer
A certificateThe admissible class is empty, or the optimum is not attainedthe case, the boundary, and the proof

The first two shapes separate uniqueness of the answer-class from whether the problem singles out a representative. A selector supplied later does not correct the earlier answer; it defines a richer problem whose complete answer is the compressed one (Thm 3.2).

F = F̂ ∘ q dependence on grounds
An answer rule depends only on the problem precisely when it is constant on every class of presentations that q identifies. Labels, coordinates, enumeration order, and temporary gauges are exactly what q forgets.
7 slots in a problem
Π = (ℒ, ℋ, 𝒞, ⊢, 𝒬, 𝒪, ∼): representation language, candidate domain, premises and constraints, consequence relation, question, required form of answer, and the rule for when two outputs count as the same answer.
4 shapes of a complete result
One equivalence class with a representative fixed by every symmetry; one orbit whose symmetry fixes the class while leaving its representatives equivalent; more than one inequivalent answer, retained as the family; an empty admissible class or unattained optimum, with its certificate.
3 affirmative forms
Non-Smuggling stated positively. Content fidelity: each proposition tracks the support its premises carry. Selection fidelity: each point choice tracks a selector the problem carries. Calculus fidelity: each operation tracks the structure the problem supplies.

Def 3.4 · Cor 3.3 · §11.4 · instrument

The structure present fixes the resolution of the result: a point, an orbit, a wider family, or an empty requested class certified by proof.
Intelligent Epistemology, §3

Verdict The factorisation is the whole engine. Everything from probability to the treatment of Gettier cases is this one equation applied where a domain has named its objects. Completeness comes with it: a report presented as complete preserves every satisfying answer, every distinction the equivalence leaves standing, and every unresolved family. Compression is legitimate only when the problem itself makes the compressed distinction irrelevant.

From the paper
An answer follows from a problem exactly when every distinction that changes the answer survives in the problem.

§3 · displayed after Theorem 3.1

The paper sets this line off on its own between the theorem and its discussion. It is the theorem in English. The rest of the argument keeps returning to it.

The proof idea · Thm 3.1argued

Dependence on grounds

HypothesesA surjection q from working presentations onto problems, and an answer rule F assigning to each presentation an answer in the answer space of its problem.

  1. Suppose F factors as F̂ ∘ q. Then two presentations of one problem have the same image under q and therefore the same answer.
  2. Conversely, suppose F is constant on every class q identifies. Define F̂(Π) as F(x) for any presentation x with q(x) = Π.
  3. Constancy on the class makes that definition independent of which x is chosen.
  4. Surjectivity of q gives existence for every problem, and the construction gives uniqueness.
  5. The transport equation extends the same demand from alternative presentations to equivalent descriptions of the problem itself.

p. 5 · Thm 3.1 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

Part II

The Ground and Its Limits

Part II

Give me a place to stand, and I will move the earth.

Attributed to Archimedes

p. 9 · §4–§5 — the paper (PDF), opens in a new tab · The episode · Part II 03 / 16

Register — the paper's own margin letters: ·

2 paper sections · §4–§5
  1. §4The Anatomy of an Inference Episode
  2. §5The Foundation Theorem

declares L ∧ C ∧ A · the inference episode p. 9 · Thm 4.1 — the paper (PDF), opens in a new tab

The answer depends on the problem and on nothing else An empty class is a finished answer

A proof is not an event of proving

An abstract inferential relation needs rules and content. An actual episode also needs something that carries the relation out. Keeping the three apart stops a proof being confused with the act of proving it, and it lets the foundation theorem state the order of business without circularity: MU is proved first, and the method is applied afterwards, including to MU.

In plain words There is a difference between a rule of logic, the thing being reasoned about, and whoever is doing the reasoning. Mix them up and you get puzzles that are not puzzles at all. Keep them apart and an old problem — do you need a method before you can pick out good cases, or good cases before you can pick out a method? — stops being a circle.

Chapter Five, The Components — The First Principle Chapter Six, Self-Grounding — The First Principle

E is an inference episode ⟺ L(E) ∧ C(E) ∧ A(E)L · Logicthe rule that makes it validC · Contentpremises and conclusionA · Agentwhatever carries it outL ∧ C ∧ ¬Aan abstract inferential relation. No episode.L ∧ A ∧ ¬Crule execution with nothing inferred.C ∧ A ∧ ¬La change of state. No valid transition.Within the class of inference episodes, none of the three can be removed while the episode is retained.

Thm 4.1 · p. 9

Plate 3

Logic, Content, Agent. The three failures around the outside are the proof: rule and content without an operator leave an abstract relation, rule and operator without content leave execution with nothing inferred, content and operator without a rule leave a change of state.

The turn

Logic, Content, Agent

Logic is the rule structure that separates valid from invalid transitions and lets inferences compose. The domain under study fixes which consequence relation is in play. Content is the subject matter represented in premises and conclusion, together with whatever makes the former bear on the latter. Symbol manipulation can instantiate a calculus; reading it as inference about a domain requires a semantics.

The agent or operator is whatever carries the transition out in an actual episode. A person. An institution. A proof assistant. An algorithm. Anything that does it. The theorem concerns that function and says nothing about consciousness. The result stays available for machine reasoners without smuggling in a philosophy of mind.

What the structure says

Why the three cannot be pulled apart

Theorem 4.1 states an equivalence: an episode is an inference episode exactly when all three aspects are present. The forward direction reads them off the description itself, since from supplies the rule, premises and conclusions supply the content, and drawing supplies the operator.

The converse is where the work is. It proceeds by watching each separation fail in a different way. Take away the operator and an abstract relation remains, with no event. Take away the content and rule execution remains, with nothing inferred. Take away the rule and a change of state remains. A change of state is not a passage from premises to a conclusion.

Remark 4.1 keeps the scope honest. The theorem is relative to the object being analysed. An abstract inferential relation contains rules and content and needs no operator; an episode adds one. Definition 1.1 keeps consequence relations available at the relational level throughout.

What the structure says

The foundation theorem

Three clauses. The order among them is the content. Proof order: MU is proved independently of the later method, which is then applied to every subsequent proof, question, and assessment. Minimality: any proposition sufficient to ground the existence of consistent inference entails MU, so MU is the weakest claim sufficient for the task. Self-application: MU has a direct proof, while every valid assessment of it is itself an inference governed by the relation MU concerns.

Reflexive application follows the proof rather than replacing it. That sequence is what distinguishes this from a bootstrap. A theory that proved its floor by using its own method would be arguing in a circle; a theory that proves its floor first and then applies the method above it is doing ordinary mathematics.

What the structure says

Criterion and content

The problem of the criterion asks whether one must begin with reliable cases and infer a criterion, or begin with a criterion and identify reliable cases. Stated that way it looks circular. It is not. It stops looking circular once method and subject matter occupy their proper roles.

The method follows from what it is for an answer to come from its inputs: relevant information is visible, equivalent descriptions agree, and a complete report preserves the full supported answer. Premises, models, and channels provide the material being assessed. Experience tests their fit to the world. The inferential form governs the assessment, and it does not compete with the material for the same job.

Fallibilism survives intact. A criterion can govern what follows from a model while the model, the representation, or the channel stays revisable. The circle dissolves. The empirical task of improving the content does not.

3 inseparable aspects
Logic supplies the rule that makes the transition valid. Content supplies the premises and conclusion. The agent or operator carries the transition out. Within the class of inference episodes, none can be removed while the episode is retained.
3 ways the separation fails
Rule and content without an operator leave an abstract relation and no episode. Rule and operator without content leave execution with nothing inferred. Content and operator without a rule leave a change of state.
0 assumptions about consciousness
The operator may be a person, an institution, a proof assistant, an algorithm, or a machine. The theorem concerns the function and takes no position on what carries it.
Experience tests their fit to the world, while the inferential form governs the assessment.
Intelligent Epistemology, §5.1

Verdict The order matters more than it looks. MU has a direct proof that uses no later machinery. Applying the method to MU afterwards is therefore not question-begging. Minimality does the rest. Anything strong enough to ground the existence of consistent inference already entails MU. That makes MU the weakest claim sufficient for the task.

From the paper
The problem of the criterion untangles once method and subject matter occupy their proper roles.

§5.1 · Criterion and content

One of the paper's characteristic moves. The classical problem is not defeated by a stronger argument; it is dissolved by noticing that two of its terms were being asked to do each other's work.

p. 10 · §6 — the paper (PDF), opens in a new tab · The ceilings · Part II 04 / 16

Register — the paper's own margin letters:

declares ∅ + certificate · the certificate p. 10 · Prop 6.1 — the paper (PDF), opens in a new tab

A proof is not an event of proving Three cuts, and then four laws

An empty class is a finished answer

Gödel, Turing, Tarski, Fitch, Arrow, Goodman, No Free Lunch, and algorithmic induction are usually heard as warnings about the reach of reason. The paper reads them as achievements of the same kind: each fixes a domain and classifies exactly what that domain permits. Set beside Shannon's positive classification, they display the principal shapes a full result can take.

In plain words The famous impossibility theorems are not bad news about thinking. They are finished pieces of thinking. Each one takes a precisely stated question and returns the complete answer. Sometimes that answer is that nothing satisfies the question as asked. That is a result. It comes with a proof. Nothing is missing from it.

Chapter Six, Self-Grounding — The First Principle Chapter Nine, The New Riddle — The First Principle

The turn

Impossibility completes a classification

Suppose a problem asks for an object satisfying a list of conditions. A proof that the admissible class is empty classifies every candidate in that class as inadmissible, so the empty answer space exhausts the requested domain. The proposition is short. Its consequence is not. An impossibility proof is a finished result. A report that returns one has answered the question. Nothing further is owed.

The constructive routes forward are equally explicit. There are four. Revise a condition. Restrict the domain. Enlarge the output type. Add structure. Each changes the problem, and the change is on the record.

What the structure says

Formal ceilings

Gödel's incompleteness theorems require a consequence relation, consistency conditions, and conclusions that follow from premises before they can be stated at all. MU is prior in logical role, because the incompleteness theorem is itself a valid result from its grounds. The domain-specific hypotheses belong to the Gödelian problem, and the complete result is the theory's internal proof closure while the expressible truths may extend past it.

Turing identifies an empty class of universal halting deciders within a model of computation. Tarski returns a level distinction rather than an emptiness: full truth for a sufficiently expressive object language lives in a metalanguage or a restricted hierarchy. Gödel marks a ceiling inside formal reasoning while presupposing the floor on which the ceiling is proved.

What the structure says

Knowability, aggregation, and generalisation

Fitch's paradox isolates a different limit. Suppose every truth is knowable. Suppose also that some truth p is unknown. Then p together with the fact that p is unknown is true, so universal knowability makes it possible to know that conjunction — which yields both that p is known and, by factivity of the second conjunct, that it is not. Under those modal rules, every truth being knowable means every truth is known.

The result classifies the unrestricted schema and names its exact revision points: the schema, the modal logic, or the conception of knowledge. Arrow does the same for aggregation, returning an empty non-dictatorial class for the full package and making the price of each relaxation visible.

Goodman's riddle and No Free Lunch meet at one lesson. Successful generalisation requires structure supplied by the problem. Inductive bias is that structure. Where representation, problem distribution, or invariance is unresolved, the full answer keeps the surviving family.

Algorithmic induction closes the section by making the relativity explicit. A universal machine supplies the coding against which simplicity is measured, and the invariance theorem bounds variation across suitable machines while preserving finite-scale reference dependence. Universal means universal relative to a chosen coding scheme. That qualification returns in Part V, where an adversarial choice of machine turns a theorem into a warning.

Nine classified results and the shape of the answer each returns (§6)
Nine classified results and the shape of the answer each returns (§6)
ResultWhat the problem givesForm of the complete answer
GödelAn expressive, effectively axiomatised theory with arithmetisation and the relevant consistency hypothesisan internal proof ceiling: the system's own consistency lies beyond its theorem closure
TuringA model of effective computation and the demand for a total halting deciderthe total halting-decider class is empty
TarskiAn expressive object language and the demand for an internally definable truth predicatefull truth moves to a metalanguage or a restricted object language
FitchFactive knowledge, modal closure, and universal knowabilityuniversal knowability collapses to universal knowledge
ArrowPreference orderings, three or more alternatives, and the fairness and independence packageevery aggregation rule satisfying the full package is dictatorial
GoodmanEvidence, a projective question, and an unresolved predicate or representation languagea family of projective rules, until representation and invariance are supplied
No Free LunchAn unrestricted problem class averaged uniformly over all objective functionsuniformly averaged performance is equal; problem structure creates advantage
Algorithmic inductionA coding language or universal reference machineuniversality and simplicity are fixed only relative to that reference
ShannonA source, channel, output alphabet, and the continuity and composition conditionsentropy and channel quantities in their classified information-theoretic form

Read the middle column first. Each row's ceiling exists because its problem was fully stated. The last row shows the same machinery returning a positive classification rather than an empty class.

∅ + certificate a complete classification
A proof that the admissible class is empty classifies every candidate in the defined class as inadmissible. The empty answer space exhausts the requested domain.
knowable ⇒ known Fitch's collapse
Factive knowledge, ordinary modal closure, and the claim that every truth is knowable. If some truth p is unknown, knowing p ∧ ¬Kp would yield both Kp and ¬Kp. Universal knowability collapses into universal knowledge.
4 constructive routes out
A revised condition, a restricted domain, an enlarged output type, or added structure. Arrow's theorem makes the price of each visible rather than hiding it inside a design failure.
The limitation of reasoning is still something known by reasoning.
Intelligent Epistemology, §6.5

Verdict Every ceiling in this table was proved by reasoning, inside a problem that had to be fully specified before the proof could run. That is the paper's structural point about limits: a limitation of reasoning is still something known by reasoning. The certificate that establishes it is a complete answer rather than the absence of one.

From the paper
More generally, MU places every limit theorem inside the result supported by its problem: a certified empty class can itself be the complete answer. The limitation of reasoning is still something known by reasoning.

§6.5 · Reference-relative universality

The italics are the paper's. The sentence closes Part II and sets up the whole of Part III, where the same move produces four positive classifications instead of a ceiling.

Part III

The Mathematics of Rational Belief

Part III

Twelve sections of the paper, §7–§18, in six here. The four laws take one section each; full answers folds §13–§15 and channels folds §16–§18.

A wise man proportions his belief to the evidence.

David Hume, An Enquiry Concerning Human Understanding
proved One rule, four laws The generated sequence assembling: scalar plausibility to probability, probability to relative entropy, relative entropy to the two projections, value and cost to choice. ≈ 24s

Source: Intelligent Epistemology · Thm 8.1 · Fig. 2 · film 2

Plate 4

One rule, four laws. The generated sequence assembling: scalar plausibility to probability, probability with hierarchical decomposition to relative entropy, relative entropy with a reference to starting belief and with retention to revision, value and cost to choice (Thm 8.1, Fig. 2).

The derivation in prose · Three cuts, and then four laws

Thm 8.1 · Four quantitative laws · Part III

Four clauses, and a proof one line long.

The theorem states the four arithmetic forms and defers every proof to the sections that follow: The next four sections prove items (i)–(iv). Each clause below names the form, the freedom it leaves, and the theorem that settles it.

  1. (i)

    Probability

    Determinate Boolean scalar plausibility, up to order-preserving regraduation.

    Thm 9.2 · p. 17
  2. (ii)

    Relative entropy

    Total informational change, with exact marginal–conditional decomposition, up to positive scale.

    Thm 10.5 · p. 22
  3. (iii)

    Least-assuming completion

    Relative-entropy projection from each permitted reference; maximum entropy under counting symmetry.

    Thm 11.3 · p. 25
  4. (iv)

    Retention

    Every feature compatible with the new constraints carried forward, which is KL projection.

    Thm 12.2 · p. 29

What rests on what · all 122 numbered results, remarks aside, at the end of this part

p. 13 · §7–§8 — the paper (PDF), opens in a new tab · Three distinctions · Part III 05 / 16

Register — the paper's own margin letters: ·

2 paper sections · §7–§8
  1. §7Three Distinctions
  2. §8The Quantitative Laws of MU

An empty class is a finished answer Probability is not chosen. It is forced.

Three cuts, and then four laws

Before any arithmetic, the paper draws three lines. Constraint against conclusion. Internal support against connection to truth. Constitutive against hypothetical. The three cut the classical problems at their joints, and §7.5 runs the cut on induction, Goodman, Gettier, and scepticism in a single paragraph each. Theorem 8.1 then announces what the next four sections prove.

In plain words Three distinctions do most of the philosophical work. What you were given versus what follows from it. Whether your reasoning was any good versus whether it happened to land on the truth. Whether a rule is part of what an activity is, or a claim made inside that activity. Get these apart and the famous problems stop overlapping.

Watch the derivation assemble · film 2

Chapter Four, Epistemic Zero — The First Principle Chapter Five, The Components — The First Principle Chapter Seven, Hume’s Ghost — The First Principle Chapter Eight, The Guillotine — The First Principle Chapter Ten, The Best Explanation — The First Principle Chapter Seventeen, The Ground Leads Somewhere — The First Principle Epilogue, One Foundation — The First Principle

What the grounds paid for

Declare what the problem contains. Then make a claim about it. The three lamps are the same requirement seen three ways. Every one of them compares the claim to an answer recomputed from the declarations rather than to an answer key.

the problem declares
and the answer claims to
all three fidelities hold · 3 supported classes, 3 reported · nothing smuggled
the candidate domain ℋ · six states, three atoms each H₁ p₂ = 0.30 H₂ p₂ = 0.30 H₃ p₂ = 0.30 H₄ p₂ = 0.10 H₅ p₂ = 0.60 H₆ p₂ = 0.60
  • content fidelity holds The report covers all 3 classes the grounds support, so nothing is asserted beyond what the premises carry.
  • selection fidelity holds No point choice was made, so there is no selector to track.
  • calculus fidelity holds No operation was used beyond reading the surviving set.
  • classes supported 3 by the grounds as declared
  • classes reported 3 by the claim as made
  • distinctions dropped 0 supported, and not reported
  • what the claim returns 3 classes: H₁ · H₂ · H₃ computed on the survivors, not looked up

Nothing smuggled. Every proposition the report makes is carried by the declared grounds, no point was chosen that the problem does not fix, and no operation was used whose structure the problem does not supply.

at machine scale §28.3 gives this failure its other name. A generative system that reports content its constraints never paid for instantiates exactly what Non-Smuggling forbids. The standard applies to machine reporters unchanged. Hallucination is this violation under thin constraints.

None of the three lamps is a test of effort or of good faith. Each one asks the same question about a different part of the report: does this piece of the answer trace back to something the problem contains? Declaring the missing structure turns every red lamp green without changing a single number, which is the point. The problem is not that the answer was wrong. It is that it was not paid for.

proved the three affirmative forms are Corollary 3.5; the licensing of a compression by a declared selector is Theorem 3.2 with Corollary 3.3, and the dependence on grounds is Theorem 3.1 (§3). The machine-scale reading is §28.3. argued the six candidate states, the stated constraint, the order, and the conjunction A ∩ B are display choices. The audit itself declares nothing: the honest answer is recomputed from the grounds on every click. The lamps compare the claim to that rather than to a stored verdict.

Source: Intelligent Epistemology §3 · §7.5 · §28.3 · Cor 3.5 · Thm 3.1 · Thm 3.2 · Cor 3.3

Plate 5

The three fidelities of Corollary 3.5, run as a check. Move the stated grounds on the left and watch each lamp: content fidelity fails when a proposition claims support its premises never carried, selection fidelity when a point is chosen with no selector in the problem, calculus fidelity when an operation uses structure the problem never supplied.

The turn

Constraint against conclusion

A constraint is part of what the problem gives. A conclusion is what those constraints support under the chosen consequence relation. Non-Smuggling governs the passage from the first to the second. Both directions of error are live.

Constraints include more than observations. Far more. Logical relations, symmetries, apparatus, form of answer, and equivalence all qualify. Treating one of them as hidden background makes the problem look more determinate than it is. Treating a conclusion as an input makes a derivation circular. The constraint language of Definition 7.2 lists the five forms the paper uses and requires anything further to be declared before it may affect an answer.

What the structure says

Internal support against connection to truth

The internal question is whether an answer follows from the agent's grounds. The external question is whether those grounds represent the world and arrive through channels that reliably track the relevant truth. They are separate coordinates. Every classical case that trades on the gap between them is exploiting that separation.

A false model can support impeccable updating. A lucky guess can be true without justification. Neither is odd. Neither observation is paradoxical once the two coordinates are drawn apart. Part III's channel and convergence sections classify exactly when internal updating tracks the world, and Part IV's treatment of Gettier is this distinction with a third coordinate added.

What the structure says

Constitutive against hypothetical

A condition is constitutive when it belongs to what the practice is — that an inferential conclusion depend on its grounds, for instance. A claim is hypothetical when it is one proposition assessed within that practice. Keeping the levels distinct is what stops the regress demands from biting.

A model's connection between past and future is a substantive question with an empirical answer. An audit of all inference is itself an inference and therefore stays inside the same constitutive form. That asymmetry is the diagnosis, and §7.5 applies it four times over: induction separates constitutive support from projective hypotheses; Goodman separates constraints from representation; Gettier separates internal support from the external route; scepticism separates global denial from local channel doubt.

The turn

What the four laws claim

Each quantitative domain names an object: plausibility, informational change, starting belief, revision, choice, or learning from a channel. MU generates the laws of each by preserving every dependence the grounds carry, every distinction the state carries, and every equivalence the descriptions carry. Four recurring movements do the work — dependence, retention, refinement, completion — and their arithmetic forms are probability, relative entropy, maximum entropy, and KL projection.

Theorem 8.1 states all four and proves none of them. Its proof is one line long and points forward. The next four sections do the work. The theorem's real content is the uniqueness claim and the exact freedom it leaves. Probability is fixed up to order-preserving regraduation. Relative entropy is fixed up to positive scale. Nothing else is free at all.

The paper is careful about what a uniqueness result means here. Logicians call it relative categoricity; the plainer phrase, which the paper prefers, is uniqueness within the domain. Change the logic, the representation, the reference, the value, the cost, or the output question and a new problem has been posed. It gets its own full answer. That answer may look nothing like the one before it.

4 quantitative laws
Probability for determinate Boolean scalar plausibility. Relative entropy for total informational change. Relative-entropy projection for least-assuming completion, with maximum entropy as its finite counting-symmetry form. KL projection for revision. Each is unique in its domain.
5 forms of constraint
Logical and partition relations; expectation or moment constraints; interval and inequality constraints; symmetries and invariances; likelihood or channel conditions supplied by the apparatus. Any further constraint has to be made explicit before it can affect an answer.
4 recurring movements
Dependence, retention, refinement, and completion. The four arithmetic forms are what those movements become once a domain names its object.
2 kinds of answer
A point answer occurs when one answer remains after the relevant equivalences and explicit selectors. A set-valued answer occurs when several inequivalent possibilities remain. The family is then the full answer.

Cor 3.5 · Thm 3.1 · §7.5 · instrument

A fully stated domain can admit one mathematical form. Logicians call this relative categoricity; the plainer phrase uniqueness within the domain captures the result. The domain names the object. MU fixes the form that follows from it.
Intelligent Epistemology, Part III opening

Verdict The theorem announcing the four laws proves nothing itself; it defers to the four sections that follow. What it does supply is the shape of the claim. Each law is unique in its domain, up to a stated freedom — an order-preserving regraduation, or a positive scale — and Full-Answer Preservation carries every remaining plurality as a family and every empty or unattained case as a certificate.

From the paper
Four recurring movements do the work: dependence, retention, refinement, and completion. Their arithmetic forms are probability, relative entropy, maximum entropy, and KL projection.

§8 · The Quantitative Laws of MU

The sentence that turns Part I's philosophy into Part III's mathematics. Each movement is a way of not adding and not discarding. Each arithmetic form is what that discipline becomes once a domain has said what it is about.

p. 16 · §9 — the paper (PDF), opens in a new tab · Probability · Part III 06 / 16

Register — the paper's own margin letters: ·

declares p(·∣·) · probability p. 17 · Thm 9.2 — the paper (PDF), opens in a new tab

Three cuts, and then four laws Count every change exactly once

Probability is not chosen. It is forced.

Name the object: a Boolean algebra of propositions, a total plausibility order on a rich scalar range, a scalar state complete for conjunction and disjoint union, and symmetric refinement. Three conditions. From them MU generates five clauses of a calculus, and from those five the product rule, the sum rule, and complementation follow with one freedom left over.

In plain words Nobody has to decide that beliefs obey the rules of probability. If you are grading propositions on a single number, that number behaves properly under and, or, and not, and you can always split a case into equally likely parts, then the ordinary rules of probability are the only thing left. Everything else has been ruled out.

Chapter Four, Epistemic Zero — The First Principle Chapter Seventeen, The Ground Leads Somewhere — The First Principle

Name the object, and the calculus is settledthe domain · Def 9.1what MU generates · Prop 9.1the calculus · Thm 9.2D1Boolean algebra, total order,rich scalar rangeD2a scalar state complete forconjunction and disjoint unionD3symmetric refinement, withdense realised rangesC1scalar substitutionC2retention of supportC3refinement coherenceC4Boolean coherenceC5continuityp(A∧B∣C) =p(A∣C)·p(B∣A∧C)p(A∨B∣C) = p(A∣C)+ p(B∣C) − p(A∧B∣C)p(¬A∣C) = 1 − p(A∣C)the one freedom left: an order-preserving regraduation pProbability is not chosen. Every step above is forced by the domain conditions and the dependence rule.

Def 9.1 · Prop 9.1 · Thm 9.2 · p. 17

Plate 6

The forcing chain. Three domain conditions on the left, the five clauses MU generates from them in the middle, the three rules of the calculus on the right. Each arrow is a step in the proof, and the dashed return marks the one freedom left: order-preserving regraduation.

The turn

Naming the object

Boolean and scalar name an object here, the way group names an algebraic object. A determinate Boolean plausibility domain has three conditions. First, a Boolean algebra of propositions with a total plausibility order whose scalar range is separable and operationally rich. Second, a scalar state complete for conjunction and disjoint union, meaning the scalar inputs and the Boolean relations exhaust the operation-relevant information. Third, symmetric refinement: the domain can be cut into equal parts, and the cutting behaves. Rational subdivisions exist. A common refinement of two descriptions preserves the represented event. The realised partial ranges are dense in their target intervals.

Notice what is not assumed. No betting interpretation. No scoring rule. No axioms about preference over acts. None of it. The conditions describe what the domain is. MU supplies every inferential law of its calculus.

What the structure says

Five clauses before any rule

Completeness of the scalar state plus dependence on grounds give scalar substitution: equal scalar states are interchangeable. That yields universal operations for conjunction and disjoint union. Full-answer preservation gives retention of support. Each operation is then strictly increasing on every nondegenerate argument whose order distinction the state preserves. Symmetric refinement and presentation invariance give refinement coherence.

Boolean coherence follows from logic alone: expressions that are logically identical represent one proposition and therefore carry one value, which delivers associativity, identity, zero, complementation, and distributivity in a single stroke. Density then fills every interval, so the operations are continuous, and a rectangle squeeze upgrades separate continuity to joint continuity.

None of the five mentions probability. Not once. They are what the object has to satisfy before any rule can be written. The rules are what remains once all five are in force.

What the structure says

The two representations

Conjunction first, then disjoint union. A continuous, associative, strictly increasing operation with the top as identity and the bottom as annihilator can be regraduated into multiplication. Symmetric refinement then additivises disjoint union on the same scale.

One freedom survives that pair: the scale could be any positive power. Boolean distributivity removes it. A proposition and its negation are disjoint and exhaust the algebra, so they sum to one, and inclusion–exclusion gives the general sum rule. Both constructions are set out in the technical block below.

The turn

The neighbours

Finite qualitative systems show the boundary. Kraft, Pratt, and Seidenberg exhibit a five-atom total order that no probability measure agrees with, and Scott supplies the necessary and sufficient cancellation conditions. That domain returns an empty representation class. By Proposition 6.1 the empty class is its complete answer. Scalar completeness, symmetric refinement, substitution, and common-refinement coherence are what select the probability domain instead.

Partial orders, vector-valued states, and context-sensitive operations get their own full answers under their own specifications. So does the non-Boolean case: projection lattices in Hilbert space are classified by Gleason-type theorems under their own dimensional, additive, and regularity conditions. The paper does not claim probability everywhere. It claims probability exactly where the domain is this one. The difference is the point.

Three remarks close the section by locating rival derivations rather than dismissing them. Dutch-book arguments supply an operational consequence of incoherence under a betting interpretation. Accuracy-first epistemology derives probabilism from dominance with respect to strictly proper scoring rules. Preference-first routes recover probability and utility together from axioms on choice. Each begins from different grounds and converges on overlapping quantitative forms. The paper marks one question as still open: whether accuracy dominance and dependence on grounds force one another.

Technical · §9.2 · §9.3

What this block carriesThe associative-representation argument that turns conjunction into multiplication, and the additive one that turns disjoint union into addition.

Monotone Cauchy

h(x + y) = h(x) + h(y) ⟹ h(x) = h(1)·x

A monotone additive function on the positive reals is linear. Additivity settles the positive rationals; rational sequences rising and falling to a real then squeeze the value by monotonicity.

If doubling the input doubles the output and the function never doubles back, the function is a straight line through the origin.

p. 18 · Lemma 9.4 — the paper (PDF), opens in a new tab

Multiplicative representation of conjunction

λ(F(x, y)) = λ(x) + λ(y)

g(x) = e−λ(x), g(F(x, y)) = g(x)·g(y)

Fix an interior point a and form its integer powers under the operation. They decrease to the bottom of the interval, each has a unique nth root, and the rational powers turn out order-dense. Define λ as the supremum of rationals whose power still dominates x. Order density makes λ a continuous strictly decreasing bijection, additivity of λ over the operation follows on rational powers and extends by continuity, and g = e−λ carries the operation to ordinary multiplication.

Keep applying the operation to one fixed value and you get a ruler. The ruler turns out to be a logarithmic one, so on its scale the operation is multiplication.

p. 18 · Thm 9.5 — the paper (PDF), opens in a new tab

Additive representation of disjoint union

h(x ⊕ y) = h(x) + h(y)

h(z·x) = h(z)·h(x) ⟹ h(x) = xm, m > 0

Symmetric refinement supplies, for every n, a unique equal part that n copies of exhaust the whole. The resulting rational points are order-dense and additive. Rescaling by their supremum makes disjoint union addition. Boolean distributivity then forces the rescaling to be a power, and the power is absorbed.

Cut the certain event into n equal pieces, count pieces, and you have a scale on which disjoint alternatives add. Distributivity then removes the last freedom in the scale.

p. 19 · Lemma 9.6 — the paper (PDF), opens in a new tab

Continuity from operational richness

A monotone one-variable section can only fail continuity by a jump, and a jump omits an open interval from the realised range. Density excludes that. A four-corner rectangle squeeze upgrades separate continuity to joint continuity.

If every value in between is reached somewhere, the operation cannot skip.

p. 17 · Thm 9.3 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

3 domain conditions
A Boolean algebra with a total plausibility order on a separable, operationally rich scalar range. A scalar state complete for conjunction and disjoint union. Symmetric refinement, with the realised partial ranges dense in their target intervals.
5 generated clauses
Scalar substitution, retention of support, refinement coherence, Boolean coherence, continuity. MU generates all five from the three domain conditions before any rule of probability is written down.
3 rules, forced
The product rule for conjunction, inclusion–exclusion for disjunction, and complementation to one. Probability is MU's unique scalar Boolean calculus, up to order-preserving regraduation.
5 atoms with no agreeing measure
Kraft, Pratt, and Seidenberg exhibit a five-atom total qualitative order that admits no agreeing probability measure. That neighbouring domain returns an empty representation class. Scott supplies the cancellation conditions separating the cases.
open whether the two routes force one another
Accuracy-first epistemology derives probabilism from dominance under strictly proper scoring rules. Preference-first routes recover probability and utility together from axioms on choice. The route here begins from dependence on grounds and the domain's richness conditions. The three converge on overlapping quantitative forms, and whether accuracy dominance and dependence on grounds force one another remains open.
Probability is Epistemic Zero written as arithmetic: every supported distinction is preserved and every undetermined distinction remains open.
Intelligent Epistemology, §9

Verdict The uniqueness claim is exact and worth stating carefully. It is not that one function is picked out; it is that every representation admitted by the domain is a monotone relabelling of a single calculus. Weaken the domain and the neighbouring cases are real: a five-atom qualitative order with no agreeing measure returns an empty representation class. A partial comparison returns the family of compatible representations.

From the paper
Probability is Epistemic Zero written as arithmetic: every supported distinction is preserved and every undetermined distinction remains open.

§9 · Probability

The sentence explains why the derivation needs both halves of the bilateral rule. Preservation of distinctions gives strict monotonicity; adding nothing beyond the constraints is what leaves the calculus with exactly one degree of freedom.

The proof idea · Thm 9.2argued

MU forces probability

HypothesesA determinate Boolean plausibility domain: Boolean algebra with a total order on an operationally rich scalar range, a scalar state complete for conjunction and disjoint union, and symmetric refinement with dense realised ranges.

  1. Completeness of the scalar state plus dependence on grounds make every represented operation factor through the scalar values.
  2. Full-answer preservation makes each operation strictly increasing wherever the state preserves an order distinction.
  3. Dense refinement fills every interval, so the operations are continuous.
  4. The associative-representation argument then regraduates conjunction into multiplication.
  5. Symmetric refinement builds rational parts of the certain event, which additivise disjoint union; distributivity forces the residual freedom to a power, which is absorbed.
  6. Complementation follows because a proposition and its negation are disjoint and exhaust the algebra.

p. 17 · Thm 9.2 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

p. 20 · §10 — the paper (PDF), opens in a new tab · Relative entropy · Part III 07 / 16

Register — the paper's own margin letters: ·

declares DKL · relative entropy p. 22 · Thm 10.5 — the paper (PDF), opens in a new tab

Probability is not chosen. It is forced. One divergence, two reference points

Count every change exactly once

One transition between probability states admits two descriptions: a flat joint account, and a staged account that reports a marginal change and then conditional changes on each branch. MU requires both to carry the same total. That single demand, plus the requirement that conditional branches be weighted by the revised state, forces relative entropy up to a positive scale.

In plain words You can describe a change all at once or in stages. If both descriptions are of the same change, they have to add up to the same amount. Insist on that and there is only one way to measure informational change. It is the one information theory already uses.

Interlude, The Story of Reasoning — The First Principle

The ledger

One transition, two descriptions. Sum the joint cell by cell, or stage it as a marginal move plus what each branch does afterwards. MU asks the two totals to agree, and everything else in this section follows from that demand.

flat 0.736463 · staged 0.736463 · disagreement zero to machine precision · coarse-grained 0.451816
the staged transition dashed: the reference conditional · filled: the revised one x₁ · weight 0.7500 y₁ y₂ y₃ x₂ · weight 0.2500 y₁ y₂ y₃
run it to
The account, both ways. The right-hand column is summed over all 6 joint cells with no staging; the left is built one term at a time.
  • the staged account weight conditional contributes
  • marginal move on X · D(QX‖PX) 0.049857
  • branch x₁ · D(QY|x₁‖PY|x₁) 0.7500 0.915475 0.686607
  • branch x₂ · D(QY|x₂‖PY|x₂) 0.2500 0.000000 0.000000
  • staged total 0.736463
  • flat total · summed over the 6 joint cells 0.736463

In nats the two accounts disagree by zero to machine precision. Both totals are computed from the state on screen: the flat one sums 6 cells, the staged one adds three terms, and neither is derived from the other.

Branch weights are QX, the revised state, as Lemma 10.1 requires. Switch them for the old state and the two accounts stop agreeing.
Coarse-graining, Lemma 10.4: push both states through one kernel that merges y₂ and y₃
  • D(Q‖P) 0.7365 before the kernel
  • D(QK‖PK) 0.4518 after it
  • the fall 0.2846 never negative, and here is why
  • the remainder 0.2846 the proof's own nonnegative term

The lemma decomposes one joint two ways. Through the merged cell it collapses; through the fine cells it carries a surplus that is a sum of divergences and cannot be negative. The fall and that surplus are computed separately here. In nats they differ by zero to machine precision.

The staged account reads 0.736463 nats and the flat one reads 0.736463. Coarse-graining then takes the total down to 0.451816, a fall of 0.284648, and that fall is exactly the information the merged cell no longer distinguishes.

The chain rule is not a convenience. It is what picks relative entropy out of the wider family of monotone divergences: drop the demand that the staged and flat descriptions balance exactly. Rényi and the rest come back in. Everything downstream of this card — the two projections, the Pythagorean split, the convergence argument — rests on the disagreement row above.

proved the accounting identity is Theorem 10.2 (I4), the final-state weight is Lemma 10.1, and the coarse-graining inequality is Lemma 10.4 (§10). argued the reference joint, the three signal cells, and the tilt direction g = (−1, 0, +1) are display choices. The identities hold on every finite informational-change domain. Every figure on this card is computed for the state shown rather than fitted to it.

The measured residuals
flat against staged
2.220e-16
the fall against the surplus
1.665e-16

Both totals are summed from the state on screen and neither is derived from the other. A figure of 0 means the two landed on the same double. Anything else is the width of one. Where a line above reports how far two accounts disagree, this is the figure it stands for.

Source: Intelligent Epistemology §10 · Thm 10.2 · Lemma 10.1 · Lemma 10.4 · Cor 10.6

Plate 7

The chain rule as a ledger. Drag a conditional branch and watch the staged account fill: a marginal term, then one conditional term per branch weighted by the revised state. The flat total on the right has to match. Push both states through the coarse-graining and the total can only fall (Thm 10.2, Lemma 10.4).

The turn

The object, and the direction

An informational-change domain represents a transition by a nonnegative scalar magnitude with the unchanged transition as its identity. It admits independent product composition, disjoint branch composition, symmetric refinement, and equivalent flat and hierarchical descriptions. Its realised magnitudes are rich enough for the continuous representation of those compositions.

Before anything else the direction has to be fixed. Lemma 10.1 fixes it. A conditional change of magnitude d on a final-state branch of probability q contributes exactly qd to the total. The proof refines the final-state partition into equal copies, uses symmetry and branch additivity on rationals, and extends by continuity. The consequence is that conditional changes are averaged with the revised marginal rather than the old one, and reversing the transition exchanges both the arguments and the branch weights.

What the structure says

Six properties, and the one that matters

The identity law and full-answer preservation give nonnegativity with strict zero only at agreement. Dependence on grounds gives invariance under common relabelling, since a relabelling preserves the transition it describes. Independent components compose by a universal scalar operation, and the associative-representation argument that ran the probability derivation now supplies a scale on which that composition is addition.

Then the chain rule arrives. A joint transition has a flat description and a staged one, and both represent the same transition, so dependence on grounds forces them equal. With final-state weighting already fixed, the equality is the exact marginal–conditional decomposition. Information monotonicity is a corollary rather than an axiom: decompose one joint two ways and the inequality falls out.

What the structure says

Pinning the constant

The characterisation proof is a chain of three moves. Write c(n) for the divergence of a point mass from the uniform on n outcomes. Splitting n into blocks makes the chain rule give additivity across products. A carefully built stochastic kernel maps the uniform on n+1 outcomes onto the uniform on n while fixing the point mass, so monotonicity gives c(n+1) ≥ c(n). Additive and monotone means logarithmic. That is the first move.

Next, refinement leaves the value alone, padding with commonly empty cells changes nothing, and a uniform nested inside a larger uniform contributes the logarithm of the ratio. Finally, build a joint whose rows are uniform blocks sized so that the first argument nests inside the second, evaluate it through the row variable and through the permuted whole, and equate. The KL formula appears on strictly positive rational pairs, and continuity extends it everywhere.

One positive constant survives. It is a unit rather than a law. Corollary 10.6 then extends the formula across the support boundary by lower semicontinuity, which is where the infinity comes from.

The turn

One divergence, two uses

The section closes with a remark that organises the next two. Prior completion and updating use one classified divergence with different reference objects. The first reference is a measure permitted by the representation and apparatus. The second is the current state of belief. The same divergence governs both maps.

That is not an economy of notation. It is the reason the two projections have the same theory: least-assuming completion and retention are the same instruction pointed at different objects. Carry forward every feature compatible with the relevant constraints. Change the rest.

Technical · §10

What this block carriesThe characterisation proof: uniform divergences grow logarithmically, refinement leaves the value alone, and two evaluations of one nested pair pin the constant.

Uniform divergences and logarithmic growth

c(mk) = c(m) + c(k)

c(n) = κ·log n, κ = c(2) > 0

Write c(n) for the divergence of a point mass from the uniform on n outcomes. Splitting n = mk into m blocks of k makes the chain rule give c(mk) = c(m) + c(k). A carefully built stochastic kernel maps the uniform on n+1 outcomes onto the uniform on n while fixing the point mass, so monotonicity gives c(n+1) ≥ c(n). Additive and monotone forces logarithmic.

p. 22 · Thm 10.5 — the paper (PDF), opens in a new tab

Refinement, padding, and nesting

D(Ū_s ‖ Ū_t) = κ·log(t/s), s ∣ t

Spreading each cell uniformly over a block leaves the divergence unchanged, because the conditional terms are uniform against uniform and vanish. Adjoining cells both states leave empty changes nothing. A uniform on s cells nested inside a uniform on t contributes κ·log(t/s).

p. 22 · Thm 10.5 — the paper (PDF), opens in a new tab

Two evaluations of one pair

D(Q‖P) = κ·Σᵢ Qᵢ·log(aᵢ / bᵢ) = κ·DKL(Q‖P)

Build a joint whose rows are uniform blocks of sizes chosen so that the first argument nests inside the second. Evaluate through the row variable and through the permuted whole. Equating the two gives the identity on strictly positive rational pairs, and continuity extends it.

p. 22 · Thm 10.5 — the paper (PDF), opens in a new tab

Closed-simplex extension

D(Q‖P) = +∞ when Q ⋠ P

The lower-semicontinuous envelope on a finite closed simplex. Where the revised state is dominated, approach every zero coordinate through strictly positive pairs. Where it is not, every approximating sequence carries a term diverging upward.

p. 23 · Cor 10.6 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

6 generated properties
Identity with strict retention; invariance under common relabelling; additive composition of independent components; exact marginal–conditional decomposition; contraction under every common stochastic channel; continuity on positive simplex interiors.
k > 0 the only freedom left
D(Q‖P) = k · DKL(Q‖P) on strictly positive finite states. One positive constant, which is a choice of unit rather than a choice of law.
D(QK‖PK) ≤ D(Q‖P) coarse-graining never adds
Push both states through the same stochastic kernel and the measured change can only fall. The inequality is derived from exact accounting rather than assumed alongside it.
+∞ when Q is not dominated by P
The lower-semicontinuous envelope of the positive-simplex formula. Where the revised state puts mass on a cell the old state leaves empty, the accounting does not close.

Thm 10.2 · Lemma 10.4 · instrument

Relative entropy is Epistemic Zero's ledger: every change is counted once, and every equivalent description balances to the same total.
Intelligent Epistemology, §10

Verdict Exact conditional decomposition is what selects KL from the wider family of monotone divergences. Weaken the chain rule to monotonicity alone and the Rényi divergences walk back in. The paper's own price list in §18 makes the trade explicit: relative entropy costs you the exact chain rule and nothing else.

From the paper
Relative entropy is Epistemic Zero's ledger: every change is counted once, and every equivalent description balances to the same total.

§10 · Relative Entropy

Counted once is the operative phrase. The uniqueness of KL is not a claim about convenience or tradition; it is what survives when double-counting and under-counting are both ruled out.

The proof idea · Thm 10.5argued

MU forces relative entropy

HypothesesA finite informational-change domain: a nonnegative magnitude for each transition, independent product composition, disjoint branch composition, symmetric refinement, and equivalent flat and hierarchical descriptions.

  1. Final-state branch weighting fixes the direction: a conditional change on a branch of revised probability q contributes q times that change.
  2. Because one transition has both a flat and a staged description, dependence on grounds forces the exact marginal–conditional chain rule.
  3. The chain rule applied to a point mass against the uniform makes the uniform divergences additive across block splittings.
  4. A stochastic kernel carrying the uniform on n+1 outcomes to the uniform on n, while fixing the point mass, makes them monotone.
  5. Additive plus monotone gives logarithmic growth with a positive constant.
  6. Nesting one uniform inside another and evaluating a single joint two ways gives the KL formula on strictly positive rational pairs; continuity finishes it.

p. 22 · Thm 10.5 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

p. 24 · §11–§12 — the paper (PDF), opens in a new tab · The two projections · Part III 08 / 16

Register — the paper's own margin letters: ·

2 paper sections · §11–§12
  1. §11Starting Beliefs: Reference Measures and Maximum Entropy
  2. §12Updating

declares μ → P₀ → P₁ · the two projections p. 23 · Remark 10.2 — the paper (PDF), opens in a new tab

Count every change exactly once Sometimes the answer is a set, and that is the answer

One divergence, two reference points

Starting belief is MU applied before evidence arrives; updating is MU applied through time. Both are the same instruction — carry forward every feature the constraints permit — pointed at different objects. Projecting from a permitted reference gives least-assuming completion, with maximum entropy as its finite counting-symmetry form. Projecting from the current state gives KL revision.

In plain words Two questions look different and turn out to be one. Where should you start before you know anything? Where should you move when you learn something? Both answers are: go to the nearest state that satisfies the constraints, where nearest is measured by the ledger from the last section. Only the starting point changes.

Watch the derivation assemble · film 3

Chapter Four, Epistemic Zero — The First Principle Chapter Fifteen, When MU Refuses to Answer — The First Principle

One divergence, two references

The same minimisation runs twice. Completion measures from the reference the problem permits; revision measures from the state belief is already in. They land in different places, and neither landing is a choice anyone made.

completion D(P₀‖μ) = 0.0302 · revision D(P₁‖P) = 0.3864 · the two landings sit 0.0264 apart in total variation
atom 1 atom 2 atom 3 𝒞₁ · E[f] = c Q, any member of 𝒞₁ μ, the permitted reference P, the current state completion revision
  • D(P₀‖μ) 0.0302 completion, from the permitted reference
  • D(P₁‖P) 0.3864 revision, from the current state
  • H(P₀) 1.0684 the completion's entropy, in nats
  • apart 0.0264 total variation between the two landings

Theorem 11.3, both sides on the completion point: DKL(P₀‖U) reads 0.030229 summed over the three atoms, and log 3 − H(P₀) reads 0.030229 from the entropy beside it. The two disagree by zero to machine precision. Least-assuming completion and maximum entropy are one problem written in two coordinate systems. This is the row that says so.

Theorem 12.6, both sides: D(Q‖P) = D(Q‖P₁) + D(P₁‖P), for every Q in 𝒞₁
  • the split D(Q‖P) D(Q‖P₁) D(P₁‖P)
  • summed over the three atoms, each term on its own 0.411293 0.024915 0.386378

The left side reads 0.411293. The two right-hand terms add to 0.411293. They disagree by zero to machine precision. Move Q anywhere along the family and the split holds. The angle it names is informational, not the angle the drawing appears to make.

Corollary 12.8: alternate between 𝒞₁ and a second family, and the walk converges on their intersection
steps to the tolerance: stopped by: not run

Not run yet. The walk starts at the current state, projects onto 𝒞₁, projects that onto 𝒞₂, and repeats. Each step is a KL projection. The corollary says the sequence reaches the projection onto the intersection.

Remark 10.2 is one sentence and it decides a great deal. Completion asks what the least-assuming state compatible with the constraints is, measured from what the problem permits before belief arrives. Revision asks what the least-assuming state compatible with the constraints is, measured from where belief already stands. Feed the same constraint to both and they part company, which is why a prior and an update are different objects governed by one law.

proved completion is Theorem 11.1 with Theorem 11.3, revision is Theorem 12.1 with Theorem 12.2, the split is Theorem 12.6, and the alternation is Corollary 12.8 with Lemma 12.7 (§11–§12). argued three atoms, the observable f = (0, 1, 2), and the second family p₂ = d are display choices; the results hold on any finite outcome space with a linear family carrying a strictly positive point. Projections are solved by bisection on the exponential-family parameter to a stated tolerance, and the alternating run carries a named clamp at 200 steps, reported whenever it is what stopped the walk.

The measured residuals
divergence against entropy
3.331e-16
the Pythagorean split
2.220e-16

Each of these is one quantity computed two ways from the state on screen, with neither route derived from the other. A figure of 0 means the two routes landed on the same double. Anything else is the width of one. Where a line above reports how far two readings disagree, this is the figure it stands for. The stated tolerance the bisection stops at is 1e-12, and the alternating run reports which of the two ended it.

Source: Intelligent Epistemology §11 · §12 · Remark 10.2 · Thm 11.3 · Thm 12.2 · Thm 12.6 · Cor 12.8

Plate 8

One divergence, two reference objects. Drop a constraint set onto the simplex and watch both projections run. Completion starts from the permitted reference μ. Revision starts from the current state P. The right angle is the Pythagorean identity; the alternating path between two constraint families is Corollary 12.8 converging.

The turn

What a permitted reference is

Before evidence, the problem still supplies something: the comparison structure against which any concentration will be measured. Definition 11.1 makes it precise. A positive measure class is permitted when it belongs to an invariant assignment across the representations the problem and apparatus allow, with the transformations that preserve the statistical question carrying one presentation's class to another's. The permitted family is every such value, taken up to positive scale. No member is preferred.

MU then selects, for each permitted reference, the feasible states carrying the least informational departure from it. One reference and one attained minimiser give a point prior. Several references or several minimisers give the complete attained family. Every unattained permitted case appears separately with its attainment boundary.

What the structure says

Maximum entropy in counting coordinates

On a finite atom space with a permutation-symmetric counting reference, the relative entropy to the uniform is log N minus the Shannon entropy. The two optimisation problems therefore have identical solutions. Maximum entropy is not a separate principle. It is least-assuming completion written in counting coordinates.

With the reference fixed and the feasible set determined by attained expectation constraints, an interior solution takes exponential-family form. That form is the Euler–Lagrange equation for a strictly convex objective under affine constraints. Declare a complexity function together with a moment constraint on it. The same theorem returns a Boltzmann prior over hypotheses. That is an Occam ordering relative to the chosen code. Change the code and the case changes with it. Say which code.

What the structure says

Where the reference comes from, and where it runs out

Problem-preserving transformations determine the reference answer, and three singleton cases are named. A compact transitive symmetry with a unique normalised invariant measure. A problem that supplies the Fisher information metric and asks for its induced volume class. An affine coordinate with congruence invariance, which gives Lebesgue and, on a bounded interval, its normalised uniform case.

Noncompact symmetry is where the machinery stops short. The paper says so precisely. Translation invariance gives every unit interval the same mass, and countable additivity over a disjoint cover of the line then forces total mass zero or infinite. Scale invariance runs the same argument on dyadic intervals. The invariant class survives. The normalisation does not. A proper probability emerges only when the problem adds location, scale, truncation, likelihood, or another normalising structure.

The turn

Revision as retention

Updating applies the same rule through time. The current state carries the information already earned; new constraints purchase a specific departure from it. Revision carries forward every feature compatible with the constraints and changes exactly what joint satisfaction requires. The complete update is therefore the whole set of divergence minimisers over the feasible set.

Ordinary conditioning is the special case where the constraint sets an event's probability to one. Bayes' rule appears there as the minimiser. Jeffrey conditioning is the case where it sets that probability to some other value: the old conditionals inside and outside the event survive, and only their mixture weights move. Both are consequences here rather than additional postulates. Neither was assumed.

On a nonempty closed convex feasible set with a point of finite divergence, strict convexity gives existence and uniqueness. Alternating projections between two linear families converge to the projection onto the intersection, which the Pythagorean identity and Pinsker's inequality prove together. The telescoped identity makes the step divergences summable. Pinsker turns that into vanishing total-variation steps. Compactness supplies the limit.

Path independence closes the section. For compatible finite linear constraints, the simultaneous projection is a function of the joint feasible set. That set is invariant under the order the constraints are presented in. Alternating projection converges to it. A single sequential pass reaches the same point when the projection operators commute or preserve one another's constraint families, and the paper is careful not to claim more: nonlinear and nonconvex constraints need their own algorithms and their own convergence theorems.

Technical · §11 · §12.2

What this block carriesExponential-family form, the Pythagorean identity, Pinsker, and the alternating-projection convergence proof.

Exponential-family form

dP*/dμ (x) = exp(Σⱼ λⱼ·fⱼ(x)) / Z(λ)

The Euler–Lagrange equation for the strictly convex relative-entropy objective under affine expectation constraints. Existence and boundary cases still need the usual integrability and attainment hypotheses.

p. 25 · Thm 11.4 — the paper (PDF), opens in a new tab

Pythagorean identity

DKL(Q‖P) = DKL(Q‖Q*) + DKL(Q*‖P)

On a linear family with a strictly positive interior point, the Lagrange equations make the log-density ratio at the projection an affine function of the constraint statistics. Its expectation is therefore the same for every feasible state, and expanding the divergence through the projection splits it exactly.

p. 30 · Thm 12.6 — the paper (PDF), opens in a new tab

Pinsker inequality

DKL(P‖Q) ≥ 2·dTV(P, Q)²

Coarse-grain by the indicator of where the first state exceeds the second and apply information monotonicity. The binary case reduces to a function with value and first derivative zero at agreement and a nonnegative second derivative.

p. 30 · Lemma 12.7 — the paper (PDF), opens in a new tab

Alternating projection converges

The Pythagorean identity telescopes, so the step divergences are summable and Pinsker sends the total-variation steps to zero. Compactness supplies an accumulation point; even and odd subsequences lie in the two closed families and their asymptotic equality puts the limit in the intersection. Strict convexity makes it unique, so the whole sequence converges.

p. 30 · Cor 12.8 — the paper (PDF), opens in a new tab

Noncompact symmetry and the normalisation obstruction

dx under translation · dx/x under rescaling

Translation invariance gives every unit interval the same mass. Countable additivity over a disjoint cover of the line then forces total mass zero or infinite. Scale invariance runs the same argument on dyadic intervals. The invariant class survives; the normalisation does not.

p. 26 · Prop 11.7 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

Four worked examples, one per shape (§11.4)
Four worked examples, one per shape (§11.4)
CaseThe structure statedThe complete answer
A pointSix outcomes with the full permutation symmetry of a fair dieuniform, entropy log 6; the symmetry route and the projection route meet at the same point
A familyTwo invariant references left standing, with neither preferred by the problemthe indexed family of reference-and-minimiser pairs; selecting one would add a preference the problem never supplied
An improper classTranslation invariance on the real linethe sigma-finite Lebesgue class together with the normalisation obstruction; further structure restores propriety
An unattained optimumTwo atoms, uniform reference, and the open constraint p₁ > ½the case, the infimum 0, and the boundary point (½, ½) that the constraint excludes

The paper closes each of the four by exhibiting it rather than describing it. That is the section's method in miniature: the shapes of Definition 3.4 are not a taxonomy imposed from outside, they are what the machinery produces when it is run.

argmin DKL(P‖U) = argmax H(P) on a finite atom space
With a permutation-symmetric counting reference, DKL(P‖U) = log N − H(P). Minimising departure from counting and maximising Shannon entropy are the same problem in different coordinates.
argminQ ∈ 𝒞 D(Q‖P) MU's retention law
Revision carries forward every feature of the current state compatible with the new constraints. It changes exactly what joint satisfaction requires. Every minimiser is retained; every unattained optimum is reported.
D(Q‖P) = D(Q‖Q*) + D(Q*‖P) the right angle
For a linear family on a finite outcome space with a strictly positive interior point, the projection splits the divergence exactly. Alternating projections between two such families converge to the projection onto their intersection.
0 the infimum, attained by nobody
On two atoms with uniform reference, minimise over the open case p₁ > ½. The infimum is 0, approached as p₁ falls to ½. The complete answer is the case, the infimum, and the excluded boundary point (½, ½).

Remark 10.2 · Thm 11.3 · Thm 12.2 · Thm 12.6 · instrument

Maximum entropy is Epistemic Zero before evidence: the selected state carries the least concentration compatible with the constraints and the reference.
Intelligent Epistemology, §11

Verdict The retention half is doing real work. A classified divergence fixes a geometry; without the instruction to change no more than the constraints require, any update policy could be paired with it. §18 lists that exact pairing as one of the rivals the retention rule excludes. Where several references or several minimisers survive, the answer is the family — and where the optimum is not attained, the answer is the case, the infimum, and the boundary.

From the paper
KL projection is Epistemic Zero through time: each new constraint changes exactly what it touches and carries every compatible feature forward.

§12 · Updating

Set this beside the maximum-entropy line from §11 and the pair is the whole of Part III's middle. Same instruction, same divergence, different starting object.

The proof idea · Thm 12.1argued

MU's retention law and KL projection

HypothesesA current state, a feasible constraint set, and a case that represents informational change by the classified divergence.

  1. New constraints purchase a departure from the current state, and the divergence orders feasible revisions by the size of that departure.
  2. A revision with strictly greater divergence introduces change the constraints did not require, while a feasible revision with less is available.
  3. MU therefore selects the minimisers, and full-answer preservation returns the whole minimiser set rather than one member of it.
  4. On a nonempty closed convex feasible set containing a point of finite divergence, strict convexity gives existence and uniqueness.
  5. Standard conditioning is the special case where the constraint sets the probability of an event to one.
  6. Jeffrey conditioning is the case where it sets that probability to q: the old conditionals inside and outside the event survive, and only their mixture weights move.

p. 28 · Thm 12.1 — the paper (PDF), opens in a new tabp. 29 · Thm 12.2 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

defined Every answer shape Mass on a simplex resolving four ways: a point fixed by symmetry, an orbit its symmetry cannot separate, a family nothing selects among, and a boundary the constraints never reach. ≈ 22s

Source: Intelligent Epistemology · Def 3.4 · §11.4 · film 3

Plate 9

Every answer shape. Mass on a simplex resolving four ways: a point fixed by every symmetry, an orbit its symmetry cannot separate, a family nothing in the problem selects among, and a boundary the constraints approach without reaching. The four settings are the paper's own worked examples (Def. 3.4, §11.4).

What follows when the answer is a set · Sometimes the answer is a set, and that is the answer

p. 31 · §13–§15 — the paper (PDF), opens in a new tab · Full answers · Part III 09 / 16

Register — the paper's own margin letters: ·

3 paper sections · §13–§15
  1. §13Full Answers Can Be Sets·
  2. §14From Belief to Choice
  3. §15Summary of MU's Quantitative Forms·

declares 𝒦 · the credal family p. 28 · Def 11.4 — the paper (PDF), opens in a new tab

One divergence, two reference points The kernel is where the world gets in

Sometimes the answer is a set, and that is the answer

When several inequivalent completions survive, the credal family is the result. A question can demand one number while its grounds determine an interval. The paper treats that as a finding rather than a failure, then shows what has to be added before a point appears: a value, a cost, a selector, or a protocol. Each addition is a new input on the record.

In plain words Some questions do not have one right answer given what you know. Pretending otherwise is not rigour. The honest reply is the whole set of answers your evidence permits, together with the boundary of what it fixes. If you need one number to act on, you have to supply something more. Say what.

Chapter Fifteen, When MU Refuses to Answer — The First Principle Chapter Twenty, Machines That Reason — The First Principle

Value, cost, and what survives

Three inputs and no others: what the problem permits, what it rewards, and what departing costs. Push value as hard as you like at the state the reference left empty. Then widen the family and watch how many actions the answer honestly contains.

value V, from action
τ = 1.00 · x₅ holds 0.000000 · undominated actions: 1 of 4
zero reference mass ghost: the reference μ · filled: the tilted state ρ* x₁ V = 1.0 0.179120 x₂ V = 2.0 0.405748 x₃ V = 2.0 0.324598 x₄ V = 0.5 0.090535 x₅ V = 3.0 0.000000
run τ to At τ = 1.00 the tilted state sits 0.1693 nats from the reference, and x₅ holds 0.000000 of the mass.
  • τ log Z 1.5157 the value at the maximum, Thm 14.1
  • Eρ*[V] 1.6851 value the tilted state earns
  • τ D(ρ*‖μ) 0.1693 what it pays to depart
  • ρ*(x₅) 0.000000 the state the reference left empty
Theorem 14.1, both sides: Eρ[V] − τD(ρ‖μ) = τ log Z − τD(ρ‖ρ*)
  • candidate state left side τD(ρ‖ρ*) what it gives up
  • ρ = ρ*, the tilted state 1.515729 0.000000 sits on the tilt, so it scores τ log Z and no state scores more
  • ρ = μ, stay at the reference 1.325000 0.190729 earns Eμ[V], pays nothing to depart
  • ρ = all mass on the best supported state earns the largest V it can reach, pays log(1/μ) to get there

The right-hand column is summed from the two states rather than read off the first, so the largest disagreement between the two sides is a measurement.

Theorem 14.4: the undominated set over the credal family, action by action
  • a₁ ties at its best across two states 1.3250 dominated
  • a₂ strong where the reference is heaviest 1.6100 survives
  • a₃ the same payoff whatever happens 1.4000 dominated
  • a₄ its best payoff sits on x₅ 0.3350 dominated

At width 0 the family holds one state. Every comparison is decided. The answer is a single action. Widen it and the comparisons stop agreeing with each other.

Nonemptiness, measured rather than repeated: the width was swept across 200 settings from 0 to 0.180 and the undominated set counted at each. Smallest count seen: 1. Times it was empty: 0.

State x₅ holds 0.000000 of the tilted mass. Action a₄ puts its largest payoff there. Neither fact helps it. Value reweights the possibilities the reference already carries. A possibility carrying zero mass has nothing to reweight.

A decision answer with several members is not an unfinished decision. It is the complete one, and it stays complete until an ambiguity rule, a reference, a loss refinement, or an institution supplies the structure that narrows it. Reporting a single action from a family that supports several is a selection with no selector in the problem.

proved the tilt and its identity are Theorem 14.1, the support statement is Corollary 14.2, the three roles are Proposition 14.3, the cost path is Proposition 14.5, and the nonempty undominated set is Theorem 14.4 (§14). argued the five states, the four utility vectors, and the shape of the credal family are display choices, not results. Expected utility is affine in the state. Dominance over the family is decided at its extreme points. The card checks those. State x₅ is not special-cased anywhere in the arithmetic; it holds zero because the reference gives it zero and a finite exponential leaves zero alone.

The measured residual
the two sides of the tilted score

Each of these is one quantity computed two ways from the state on screen, with neither route derived from the other. A figure of 0 means the two routes landed on the same double. Anything else is the width of one. Where a line above reports how far two readings disagree, this is the figure it stands for.

Source: Intelligent Epistemology §14 · Thm 14.1 · Cor 14.2 · Prop 14.3 · Thm 14.4 · Prop 14.5

Plate 10

A credal family, a finite action set, and the two ways to get from one to the other. Move value and the cost of departure and watch the tilt run; watch the undominated set shrink and refuse to empty. The blacked-out region is Corollary 14.2: zero reference mass stays zero however hard the value pushes.

The turn

Plurality has several causes

A credal state is the full set of point-valued answers a problem permits, and it coincides with the raw feasible simplex only when the problem independently supports every feasible point. Plurality can arise from an incomplete plausibility comparison, from several permitted reference structures, or from several tied minimisers. Attainment data stay a separate part of the answer.

Sparse evidence, unresolved representation choices, several natural references, and uncertain channel models each produce a plural state. None of them is a defect in the reasoning. Each is a fact about what the grounds fix. Report it as one.

What the structure says

Uniqueness and permissivism, reconciled

The long dispute between rational uniqueness and epistemic permissivism turns out to concern two different levels. Identical problems determine the same complete permission set. It may contain one member or many. Uniqueness governs the answer as a whole. Permissivism describes its members. Both were right.

Agents sharing a credal family may still act differently, because utilities, losses, and resources belong to the downstream decision problem. Different evidence, representation, reference families, channel models, or computational access define different belief problems in the first place. Full identity of every answer-relevant input yields identity of the full permission set. That is the only uniqueness claim the paper makes.

What the structure says

When one number is required anyway

A reporting or decision protocol may demand a point while the belief state remains a set. Resolute completion supplies one: given a reference with finite divergence to some member, take the divergence-minimising member of the family and hold it fixed until new information changes the problem. On a closed convex family with a finite-divergence point, strict convexity makes it unique. Each permitted reference yields its own resolute case. A one-number protocol becomes determinate only when it also supplies a reference or a selector.

The alternative route chooses an action directly from the credal set. Credal dominance is transitive and irreflexive, so the undominated set is never empty. Every action outside it is worse under every state the problem permits. A singleton fixes the action outright. Several survivors are the complete decision answer until an ambiguity rule, reference, loss refinement, or institutional protocol arrives.

The turn

Choice, and the boundary it inherits

Once a value functional and a cost of departure are supplied, the geometry from §10 fixes the choice distribution outright. Expected value minus the departure cost is maximised at exactly one state. The proof is an exact rewriting rather than a bound. The objective equals a constant minus the divergence to the optimum, so nonnegativity of divergence reads off both the maximiser and its uniqueness.

The corollary that follows is the sharpest line in Part III. The chosen state is equivalent to its reference in both directions. Anything the reference gives zero mass keeps zero mass. A hard exclusion can be represented by restricting the reference, which removes support. Nothing can create it. Ever. Values reweight the possibilities the reference already carries, and a truth, action, or population outside the represented measure class enters only through an enlarged representation.

The two ends of the cost path are worth holding on to. As the cost of departure grows, the chosen state returns to the reference. As it falls to zero, mass concentrates on the maximisers of value inside the support, weighted in proportion to their reference masses. Part V uses exactly this when it prices alignment.

Section 15 gathers the sequence in one plate. Scalar plausibility yields probability. Probability with hierarchical decomposition yields relative entropy. Relative entropy with a permitted reference yields starting belief; with retention it yields revision; with value and cost it yields choice. Each arrow names the structure a domain supplies. Each box is the form MU forces once that structure is present.

Technical · §14

What this block carriesThe exact rewriting behind the choice theorem, and the geometry it leaves on a smooth state space.

The identity behind the maximiser

Eρ[V] − τ·DKL(ρ‖μ) = τ·log Z − τ·DKL(ρ‖ρ*)

Because the optimal density is strictly positive almost everywhere, the log-density ratio to it splits into three terms. Integrating against any admissible state gives an exact identity, not a bound. Nonnegativity of relative entropy then reads off the maximiser and its uniqueness in one line.

p. 33 · Thm 14.1 — the paper (PDF), opens in a new tab

Score decomposition

∇ log(dρ*/dvol) = ∇ log(dμ/dvol) + (1/τ)·∇V

On a state space carrying a metric and volume form, with a strictly positive smooth reference density and a smooth value, the local score of the tilted state is the sum of a reference term and a value term relative to that geometry.

p. 35 · Prop 14.6 — the paper (PDF), opens in a new tab

A reversible realisation

dXt = ∇ log p(Xt)·dt + √2·dWt

Under the usual smoothness, nonexplosion, confinement, and ergodicity hypotheses, the overdamped Langevin diffusion is reversible with the tilted law as its stationary measure. Within the isotropic reversible class this is the standard relaxation; preconditioned and nonreversible processes reach the same law by other routes.

p. 35 · Remark 14.1 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

3 separate inputs
The reference μ records the comparison structure before value is applied. V records what the decision problem rewards. τ sets the cost of departing from the reference in units of value. Changing any one changes the decision problem.
dρ*/dμ = eV/τ / Z the unique maximiser
Expected value minus τ times the departure from the reference is maximised at exactly one state. The identity behind it is an exact rewriting rather than an approximation.
0 mass created by value
The choice from reference, value, and cost is equivalent to its reference measure in both directions. A hard exclusion can remove support. Nothing in the tilt can create support the reference did not already carry.
≥ 1 surviving actions, always
Credal dominance is transitive and irreflexive, so a finite action set always has a maximal element. One survivor fixes the action. Several survivors are the complete decision answer until a rule, reference, or protocol refines them.

Thm 14.1 · Cor 14.2 · Thm 14.4 · Prop 14.5 · instrument

A system that returns the answer’s true shape displays greater rational power than one that compresses every problem into a point.
Intelligent Epistemology, Remark 13.1

Verdict Refusal here is lawful rather than evasive, and it comes with content: the family, and the exact boundary of what the constraints determine. A system that returns the answer's true shape is doing more work than one compressing every problem into a point. When a decision is genuinely required, credal dominance removes every dominated action and always leaves at least one standing.

From the paper
For a human reasoner this is suspension of judgement. For a machine it is lawful refusal. A system that returns the answer's true shape displays greater rational power than one that compresses every problem into a point.

Remark 13.1 · Lawful refusal carries a complete answer

The line the machine sections in Part V are built on. A model that always produces a point is not more capable than one that reports a family; it is reporting a result its constraints never bought.

p. 36 · §16–§18 — the paper (PDF), opens in a new tab · Channels · Part III 10 / 16

Register — the paper's own margin letters: ·

3 paper sections · §16–§18
  1. §16Evidence Channels
  2. §17Convergence to Truth·
  3. §18What Each Condition Selects·

declares K(s∣h,η) · the evidence channel p. 36 · Def 16.1 — the paper (PDF), opens in a new tab

Sometimes the answer is a set, and that is the answer The trilemma has a fourth case

The kernel is where the world gets in

Everything so far has been internal. Evidence from the world arrives through stochastic kernels: perception, memory, introspection, testimony, instruments, models. They differ in signals, errors, dependencies, incentives, and hidden source conditions. Once the kernels are specified they share one update form. A single reliability number turns out to be a special summary rather than the general case.

In plain words A signal is not evidence until you say how the world produces it. The same reading can support a claim, undercut it, or say nothing at all, depending on the instrument. Write the instrument down and everything becomes calculable. Leave it out and the number in front of you means nothing in particular.

Introduction, The Question in the Rain — The First Principle Chapter One, The Training Run — The First Principle Chapter Seven, Hume’s Ghost — The First Principle Chapter Twelve, Do You Really Know? — The First Principle Chapter Seventeen, The Ground Leads Somewhere — The First Principle Chapter Eighteen, Science Derived — The First Principle Chapter Nineteen, Thinking Together — The First Principle Chapter Twenty, Machines That Reason — The First Principle Chapter Twenty-One, The Transition — The First Principle

The channel bench

Send a signal through and read what it is worth. The signal value is not the evidence. What the signal is worth is a property of the kernel that produced it. The same value arrives with three different verdicts under three permitted kernels.

S = 1 · Λ = 4.0000 under the calibrated kernel · posterior 0.5000 from a base rate of 0.20
send and read it under
P(H) before the signal, and after it prior 0.200 posterior 0.500 0 1
Proposition 19.5: one signal value, three permitted kernels, three verdicts
  • source state η P(s | H) P(s | ¬H) Λ(s) verdict
  • calibrated 0.8000 0.2000 4.0000 confirms
  • saturating 0.9000 0.6000 1.5000 confirms
  • adversarial 0.2000 0.8000 0.2500 disconfirms

The signal value is fixed and the three rows disagree about what it is worth. Nothing in the token settles the question. The answer is a property of the kernel the problem permits.

Proposition 16.3, the certificate: what the output distribution does not identify
q = 0 q = 1 r = 1 r = 0 filled: the pair · open: its twin · the curve: every pair with one observable law
  • Pr(S = 1) at (q, r) 0.380000
  • at (1 − q, 1 − r) 0.380000
  • signals observed log-likelihood difference between the pair and its twin
  • 10 exactly zero
  • 1,000 exactly zero
  • 1,000,000 exactly zero

Both pairs put 0.380000 on a positive signal. They disagree by zero to machine precision. The difference of log-likelihoods is computed at each sample size from the counts a run of that length would produce. It does not fall with data: it starts at zero. Self-agreement supplies consistency. Truth calibration needs an anchor from outside the channel.

A source can be perfectly consistent with itself, report at a stable rate, survive every internal audit, and still be inverted. The observable law is one number and two truth models fit it exactly. This is what bootstrapping and easy knowledge come to when the channel is written down: not a mistake in the reasoning, but a fact about what the reasoning has to work with.

proved the channel is Definition 16.1, the likelihood-ratio account of evidential force is Proposition 19.5, and the non-identifiability witness is Proposition 16.3 (§16, §19.4). argued the sensitivity, specificity, base rate, and bias settings are display choices: the paper states the identities and names no numbers for them. Every figure above is computed for the kernel on screen. The log-likelihood row uses the expected counts a run of that length would produce, and the difference it reports is zero because the two pairs induce one Bernoulli parameter, not because the row rounds.

The measured residual
the two pairs on a positive signal
5.551e-17

Each of these is one quantity computed two ways from the state on screen, with neither route derived from the other. A figure of 0 means the two routes landed on the same double. Anything else is the width of one. Where a line above reports how far two readings disagree, this is the figure it stands for.

Source: Intelligent Epistemology §16 · §19.4 · Def 16.1 · Prop 16.3 · Prop 19.5

Plate 11

Set sensitivity, specificity, base rate, and the source state, then send one signal through. The same signal value confirms, disconfirms, or says nothing depending on the kernel. The lower panel is the identifiability certificate: two truth-prevalence and accuracy pairs producing one observable law, with no signal able to separate them (Def. 16.1, Prop. 16.3, Prop. 19.5).

The turn

One object for six kinds of source

An evidence channel is a hypothesis space, a signal space, a space of source or apparatus states, and a stochastic kernel joining them. Observing a signal gives the posterior by Bayes' rule, subject to whatever point, family, or nonexistence structure the prior already carried. If several channel models remain possible, each is carried forward.

Perception and instruments are kernels from world states to registrations. Illusion, masking, calibration drift, and limited resolution are modelled by the same object. Memory is a temporally extended reconstruction channel whose latent state can include retention, interference, present cues, and rehearsal. Introspection is fallible in its classification even where the occurrence is immediate. Testimony moves odds by a likelihood ratio in which competence and honesty contribute differently. Adversarial sources need a kernel conditional on goals, incentives, and adaptation to the evaluator. A priori consequence stays internal throughout. It always was.

Unknown channel parameters are ordinary hypotheses. Calibration data update their joint posterior alongside the world hypotheses, and integrating over that posterior gives a predictive kernel richer than any raw sample proportion.

What the structure says

What a source cannot tell you about itself

Here is the section's hardest result. Its proof is two lines. In a binary symmetric channel with truth prevalence q and accuracy r, the observable rate of positive signals is qr + (1−q)(1−r). Many pairs give the same value. A pair and its double inversion give it exactly. Observed signals determine their marginal law. Nothing more.

Truth-conditional calibration additionally requires a decomposition into latent truths and a conditional kernel. Later-verified outcomes, independent controls, another calibrated route, or sufficiently restrictive structural assumptions can supply it. Repeated self-agreement cannot do it. Bootstrapping and easy-knowledge failures are cases of this non-identifiability. The diagnosis is structural rather than a complaint about circular reasoning.

What the structure says

Closure splits in two

If a body of grounds internally licenses a proposition and the consequence relation carries that proposition to another, the grounds internally license the second. The proof is one sentence. Closure is composition in the chosen consequence relation.

Route robustness is a different property. It does not travel automatically. A channel may reliably distinguish an ordinary hand-present case from an ordinary hand-absent one while failing to distinguish either from an adversarial simulation. The entailment is internal; the additional discrimination demand is a property of the kernel. Nothing in logical closure supplies it. Nothing could.

That single split is what Part IV uses on Gettier, on closure debates, on contextualism, and on scepticism. Two theorems, four pages apart. Between them they do most of the work of a literature.

The turn

Convergence, and its price

When the channel is specified, the truth is represented and supported, observations follow the sampling law, and false rivals are distinguishable, posterior mass concentrates on the represented truth exponentially fast. The proof bounds the marginal likelihood below on a small KL neighbourhood of the truth and bounds the far set above using Pinsker, then combines.

Every one of those four conditions is world-facing. Realisability holds when the true process is represented in the hypothesis space. Model enlargement is the only remedy for its failure. Distinguishability fails for observationally equivalent hypotheses, and if two hypotheses make identical predictions for all possible observations, no amount of evidence separates them. That is a limit of empirical inquiry rather than of the update rule.

The paper adds the finite-sample counterpart rather than leaving the asymptotic claim standing alone. Uniform convergence over a hypothesis class holds exactly when the class has bounded capacity, and probably-approximately-correct bounds convert capacity into sample sizes. Conformal methods trade the model for exchangeability and return finite-sample coverage from that single stated condition. The accounting is the same in each case: a guarantee holds inside the structure a method states, and the stated structure carries the price.

Technical · §17

What this block carriesThe posterior-convergence proof, bounded below and above, and the law of large numbers it rests on.

Law of large numbers, bounded case

Pr(|Z̄ₙ − μ| > ε) ≤ 2·exp(−n·ε² / (2M²))

Hoeffding's bound follows from Markov's inequality applied to the exponential moment together with the elementary sub-Gaussian estimate. The bound is summable in n, so Borel–Cantelli gives almost-sure convergence.

p. 39 · Lemma 17.1 — the paper (PDF), opens in a new tab

The denominator

Choose a closed KL neighbourhood of the truth with positive prior mass. On a finite alphabet the log-likelihood ratios are uniformly bounded and continuous there, and the empirical frequencies converge, so the normalised log-likelihood converges uniformly on the neighbourhood. The marginal likelihood is eventually at least the prior mass times an exponentially small factor.

p. 40 · Thm 17.2 — the paper (PDF), opens in a new tab

The numerator

posterior(far set) ≤ Π(B)⁻¹·exp(−n·δ²/4 + 2n·ε)

Once the empirical distribution is close to the truth, any hypothesis far in total variation is far from the empirical distribution too, and Pinsker converts that distance into a divergence gap. The far set therefore loses likelihood at an exponential rate.

p. 40 · Thm 17.2 — the paper (PDF), opens in a new tab

Support is a separate requirement

Consistency requires the represented truth to lie in the support of the completed prior. A full-support reference helps; the constrained projection decides. Common finite affine cases preserve support when the reference is strictly positive, the feasible family has an interior point, and the attained exponential-family solution is interior.

p. 39 · §17 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

What each condition selects: the rival it excludes (§18)
What each condition selects: the rival it excludes (§18)
ResultRival surviving weaker conditionsExcluding condition
Probabilityfinite qualitative orders with no agreeing measurerich refinement, scalar composition, and the corresponding coherence
Probabilityvector-valued or context-sensitive plausibility statesa one-dimensional scalar output and universal composition
Probabilitydiscontinuous associative operationsorder-density and retention of distinctions already supported
Relative entropyRényi and other monotone divergencesthe exact conditional chain rule
Prior completionone chosen reference among severala reference family fixed by the representation and symmetries
Revisiona classified divergence paired with an arbitrary update policythe retention rule
Updateone selected posterior from a minimiser familyselection fixed by the problem
Channelsagreement among correlated or strategically dependent sourcesa joint stochastic kernel and specified dependence structure
Convergencerational updating under a misspecified or non-identifiable modelrealisability, support, stable sampling, distinguishability

Every condition in Part III earns its place by excluding a named rival. Read the table as a price list: the middle column is what you get back if you decline to pay.

4 slots in a channel
The world or hypothesis space, the signal space, a space of source, apparatus, or context states, and the stochastic kernel that joins them. The source state may encode calibration, competence, incentives, selection effects, common causes, or adversarial control.
(q, r) ≡ (1−q, 1−r) the non-identifiability witness
In a binary symmetric channel with truth prevalence q and accuracy r, the observable rate is qr + (1−q)(1−r). That pair and its inversion produce the same law. Self-agreement supplies consistency data; truth calibration needs an external anchor.
4 world-facing conditions
Realisability, prior support, stable sampling, distinguishability. Given all four, posterior mass on the far set falls exponentially and concentrates on the represented truth. Drop one and the guarantee goes with it.
2 closure results, pulling apart
Internal justification is closed under the stated consequence relation, because closure is composition. Route robustness transfers only when the route also discriminates the alternatives the entailed proposition introduces.

Def 16.1 · Prop 16.3 · Prop 19.5 · instrument

A single reliability number is only a special summary. Logical consequence remains an internal relation.
Intelligent Epistemology, §16

Verdict This is the boundary between rational method and empirical success. The internal rule fixes how evidence changes belief. Channel, support, realisability, and distinguishability decide whether that process finds the truth, and all four are world-facing conditions the reasoner does not control. The convergence theorem holds when they do; §18 lists what survives when they do not.

From the paper
Here is the boundary between rational method and empirical success. The internal rule fixes how evidence changes belief; channel, support, realisability, and distinguishability determine whether that process finds the truth.

§17 · Convergence to Truth

Printed immediately after the convergence figure. It is the most load-bearing sentence in Part III. Every reconstruction of a knowledge problem in Part IV runs through it.

The proof idea · Thm 17.2argued

Posterior convergence

HypothesesObservations i.i.d. on a finite alphabet under a specified channel, with the prior giving positive mass to every KL-neighbourhood of the true per-observation distribution.

  1. Fix a small closed KL neighbourhood of the truth with positive prior mass.
  2. On a finite alphabet the log-likelihood ratios are bounded and continuous there, and empirical frequencies converge almost surely.
  3. So the marginal likelihood is eventually at least the prior mass of that neighbourhood times an exponentially small factor.
  4. Any hypothesis far in total variation is eventually far from the empirical distribution, and Pinsker turns that into a divergence gap.
  5. The far set therefore loses likelihood exponentially, and choosing the neighbourhood small enough makes the ratio vanish.
  6. For a finite identifiable class a small enough radius isolates the truth, so posterior mass on it tends to one.

p. 40 · Thm 17.2 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

The proof map

What rests on what

The paper states 122 numbered results, remarks aside, across 33 sections. Each column below is one section. Each dot is a statement, in the order that section states it. An arc runs from a statement back to the earlier one it cites by number or names by title.

The paper states 122 numbered results, remarks aside, across 33 sections. Every one of them is listed below, searchable, in the order the paper states them. Open one and it gives its text, its proof, what it rests on, and what rests on it.

37 theorems · 33 propositions · 9 corollaries · 8 lemmas · 35 definitions · 52 references

Point at a dot to light what it rests on and what rests on it. Choose one to read it.

Skip the result index

Every result

122 results

  1. Def 1.1 Inferential relation and inference episode §1
  2. Def 1.2 Consistent inference §1
  3. Def 1.3 MU §1
  4. Thm 2.1 Existence of consistent inference §2
  5. Cor 2.2 No consistent denial §2
  6. Def 3.1 Inference problem §3
  7. Def 3.2 Presentations and problems §3
  8. Thm 3.1 Dependence on grounds §3
  9. Def 3.3 Complete answer §3
  10. Thm 3.2 Full-answer preservation §3
  11. Cor 3.3 Canonical point selection §3
  12. Cor 3.4 Method and full answers §3
  13. Cor 3.5 Non-Smuggling as fidelity to grounds §3
  14. Def 3.4 Four useful shapes of a complete result §3
  15. Thm 3.6 Deductive closure §3
  16. Thm 4.1 The three aspects of an inference episode §4
  17. Thm 5.1 Foundation theorem §5
  18. Prop 6.1 Impossibility completes a classification §6
  19. Def 7.1 Constraint and conclusion §7
  20. Def 7.2 Constraint language ℒ §7
  21. Def 7.3 Internal and external §7
  22. Def 7.4 Constitutive and hypothetical §7
  23. Thm 8.1 Four quantitative laws generated by MU §8
  24. Def 8.1 Point and set-valued answers §8
  25. Def 8.2 Representation language §8
  26. Def 8.3 Hypothesis space §8
  27. Def 9.1 Determinate Boolean plausibility domain §9
  28. Prop 9.1 MU generates the probability calculus §9
  29. Thm 9.2 MU forces probability §9
  30. Thm 9.3 Continuity from operational richness §9
  31. Lemma 9.4 Monotone Cauchy §9
  32. Thm 9.5 Multiplicative representation of conjunction §9
  33. Lemma 9.6 Additive representation of disjoint union §9
  34. Def 10.1 Informational-change domain §10
  35. Lemma 10.1 Final-state branch weighting §10
  36. Thm 10.2 MU generates exact information accounting §10
  37. Prop 10.3 Relabeling invariance §10
  38. Lemma 10.4 Information monotonicity from exact accounting §10
  39. Thm 10.5 MU forces relative entropy §10
  40. Cor 10.6 Closed-simplex extension §10
  41. Def 11.1 Permitted reference measures §11
  42. Def 11.2 Projection relative to a reference §11
  43. Def 11.3 Prior answer §11
  44. Thm 11.1 MU forces least-assuming completion §11
  45. Thm 11.2 Prior completion §11
  46. Thm 11.3 MU forces maximum entropy under finite counting symmetry §11
  47. Thm 11.4 Exponential-family form §11
  48. Thm 11.5 Complexity-constrained prior §11
  49. Thm 11.6 MU fixes the reference answer §11
  50. Prop 11.7 Noncompact symmetry determines an invariant measure class §11
  51. Lemma 11.8 Interval equiplausibility §11
  52. Def 11.4 Credal set §11
  53. Thm 12.1 MU's retention law §12
  54. Thm 12.2 MU forces KL projection §12
  55. Prop 12.3 Conditions for sequential agreement §12
  56. Lemma 12.4 Convexity of the finite constraint language §12
  57. Prop 12.5 Finite-space extension §12
  58. Thm 12.6 Pythagorean identity §12
  59. Lemma 12.7 Pinsker inequality §12
  60. Cor 12.8 Projection existence, uniqueness, and alternation §12
  61. Thm 12.9 Path independence §12
  62. Thm 13.1 MU returns the credal family §13
  63. Prop 13.2 Full-answer uniqueness with plural members §13
  64. Def 13.1 Resolute completion §13
  65. Thm 13.3 Resolute completion §13
  66. Thm 14.1 Choice from reference, value, and cost §14
  67. Cor 14.2 Zero support stays zero §14
  68. Prop 14.3 Reference, value, and cost are separate inputs §14
  69. Def 14.1 Credal dominance §14
  70. Thm 14.4 Decision under imprecise belief §14
  71. Prop 14.5 The information-cost path §14
  72. Prop 14.6 Score decomposition §14
  73. Def 16.1 Evidence channel §16
  74. Prop 16.1 Channel update §16
  75. Def 16.2 Reliability terminology §16
  76. Prop 16.2 Learning a channel §16
  77. Prop 16.3 Channel identifiability §16
  78. Prop 16.4 Channel composition §16
  79. Prop 16.5 Base-rate conditional trust §16
  80. Def 16.3 Robustness of the route §16
  81. Def 16.4 Empirical knowledge proposal §16
  82. Thm 16.6 Internal justification is closed §16
  83. Thm 16.7 Route conditions govern transmission §16
  84. Lemma 17.1 Law of large numbers, bounded case §17
  85. Thm 17.2 Posterior Convergence §17
  86. Prop 19.1 Grounding result §19
  87. Prop 19.2 Rule use is irreducible to an added premise §19
  88. Prop 19.3 Finite data support a family of continuations §19
  89. Def 19.1 Higher-order problem §19
  90. Thm 19.4 Higher-order answer §19
  91. Prop 19.5 A signal gains evidential force through its channel §19
  92. Def 20.1 Deductive certainty, graded belief, and empirical knowledge §20
  93. Prop 20.1 Reference-class underdetermination §20
  94. Prop 20.2 The lottery lesson §20
  95. Prop 20.3 Finite Dutch-book coherence §20
  96. Prop 20.4 Chance–credence bridge §20
  97. Prop 20.5 Reflection under stable updating §20
  98. Prop 20.6 The logical stability of MU §20
  99. Def 20.2 Knowledge profile §20
  100. Def 20.3 Certainty §20
  101. Def 20.4 Doubt §20
  102. Thm 21.1 Conditional convergence of inductive learning §21
  103. Prop 21.2 Likelihood determines confirmation §21
  104. Prop 21.3 Confirmation follows the specified hypothesis §21
  105. Prop 21.4 Old evidence is new relational information §21
  106. Prop 21.5 Projectibility depends on representation §21
  107. Def 21.1 Projectibility conditions §21
  108. Thm 21.6 Observational equivalence §21
  109. Def 22.1 Route robustness §22
  110. Prop 22.1 Channel calibration requires an external anchor §22
  111. Prop 22.2 Disagreement as evidence §22
  112. Thm 25.1 Empirical equivalence quotient §25
  113. Cor 25.2 Maximal empirical content §25
  114. Prop 25.3 Model enlargement gives new alternatives standing §25
  115. Thm 27.1 Agreement under common priors §27
  116. Def 27.1 Common knowledge §27
  117. Thm 28.1 Causal direction requires identifying structure §28
  118. Prop 30.1 Practical reasons govern action §30
  119. Def 31.1 Three kinds of claim §31
  120. Thm 31.1 Norms of inference §31
  121. Thm 32.1 Information weakly increases optimal instrumental value §32
  122. Cor 32.2 When inquiry is instrumentally required §32

Theorem 2.1the MU theorem

Existence of consistent inference

§2 The MU Theorem · Part I · page 4

Every reflexive inferential relation whose language contains at least one consistent proposition contains a consistent inference.

The proof, as the paper gives it

Let P be a consistent proposition. Reflexivity gives the valid instance P ⊢ P. Its premise set is consistent, so it is a consistent inference and MU follows.

What does this rest on?

1 statement stands under this one. The longest chain runs 1 step to Def 1.2, where the numbered references stop.

The longest chain

  1. 1Def 1.2Consistent inference

How to read it

  • A theorem. A proposition and a corollary are the same mark, smaller and lighter.
  • An open ring is a lemma. The smallest ring is a definition.
  • A ringed dot is one of the 7 landmarks named above.
  • A citation: the paper prints the number. 27 of them.
  • A name: a statement or its proof uses, word for word, the title the paper gave an earlier one. 25 of them.

Nothing else is drawn. Prose that gestures at a result without naming or numbering it is left alone. A guess in a dependency graph is a false edge. 52 statements stand in at least one of the two relations. The other 70 are drawn quiet. A trace ends where the numbered references stop. The paper declares no axioms; that is the floor this map can show.

The 35 remarks are not here. §1.1 grades them as locating rather than proving. Neither are the 49 argument blocks, which carry no number and are cited by subsection.

Source: Intelligent Epistemology — MU and Epistemic Zero. Statements and proofs are the paper's own words, read out of the LaTeX source. The layout is a layered DAG, computed once at build.

Part IV

Classical Problems: Resolutions, Reconstructions, and Boundaries

Part IV

The register changes here. Parts I to III prove and define; from §19 the paper argues, and forty-nine of its argument blocks live in this part and the two after it.

Our knowledge can only be finite, while our ignorance must necessarily be infinite.

Karl Popper, Conjectures and Refutations

p. 43 · §19–§20 — the paper (PDF), opens in a new tab · Grounds and rules · Part IV 11 / 16

Register — the paper's own margin letters: ·

2 paper sections · §19–§20
  1. §19Foundations, Rules, and Experience
  2. §20Belief, Uncertainty, and Action

The kernel is where the world gets in A right answer can still be bad inference

The trilemma has a fourth case

The Münchhausen Trilemma (Albert, 1968) classifies inferential support as regress, circle, or unsupported stopping. MU opens a fourth: a finite direct existence proof that terminates. From that foothold the section works through Carroll's tortoise, rule-following, the myth of the given, indifference, the lottery, Dutch books, Moore's paradox, and the structure of knowledge itself.

In plain words The old trap says every justification either goes on forever, curls back on itself, or stops somewhere arbitrary. There is a fourth option nobody used: a short proof that finishes. Once you have that, a lot of famous puzzles turn out to be asking two questions at once. They come apart cleanly.

Introduction, The Question in the Rain — The First Principle Chapter Two, The First Crack — The First Principle Chapter Three, The Gallery — The First Principle Chapter Eleven, The Lottery — The First Principle Chapter Twelve, Do You Really Know? — The First Principle Chapter Fifteen, When MU Refuses to Answer — The First Principle

The four cases

The classification offers three ways for support to run out: it goes on for ever, it comes back round, or it stops on an assertion. Walk each one and read the ledger. Then walk the fourth, which the classification never enumerated.

regress · 0 moves · 1 support demand open · 0 discharged
each node is supported by the one to its right 1 1 demand open
  • moves taken 0 steps the reader has walked
  • demands open 1 support still owed
  • discharged 0 demands closed, not relocated
  • established 0 propositions beyond the premises
  • accepted unproved 0 stipulations in the walk
  • hypotheses in force 1 a support relation on propositions

The two right-hand cells measure different things. A stipulation is something the walk accepts as it goes. A hypothesis is what the case stands on before it moves: the fourth case discharges its support demand and still runs on Theorem 2.1's two.

One claim, one open support demand. Take a step and watch where the demand goes.

The denial, separately: assert ¬MU, that no consistent inference exists (Cor 2.2)
attempts: 0 consistent denials found: 0
  1. Suppose ¬MU is consistent. It is then a consistent proposition of the language.
  2. Reflexivity applies to it like any other. ¬MU ⊢ ¬MU is a valid instance.
  3. Its premise set is consistent. The instance is a consistent inference.
  4. That inference witnesses ∃I CI(I), which is the existence ¬MU denies.

Not attempted. The steps above are the whole argument; press the button and the card walks them with the counter running.

Each step discharges nothing. The support demand moves to a fresh proposition and arrives there intact, so the count of open demands is one after every move ever made. Nothing about the walker fails. The chain has no end to reach.

proved the classification and the fourth case are Proposition 19.1; the fourth case's three moves are the proof of Theorem 2.1, and the denial panel is Corollary 2.2 (§19.1, §2). The fourth case's two hypotheses are Theorem 2.1's own, taken from its statement. argued counting one hypothesis against each of the three horns reads the classification as a classification of inferential support, which is how §19.1 states it. The lane drawing, the three-node loop, and the three-step stopping walk are display lengths chosen to fit the frame. The counts they feed are computed from the walk itself. The regress carries a named display clamp at 40 moves — the walker stops there and the chain does not. located Corollary 2.2 holds in any consequence structure satisfying Theorem 2.1's hypotheses whose language contains the proposition ¬MU. A paraconsistent consequence relation defines a different theorem-generation problem with its own closure (Remark 3.3). The denial is not refuted there; it is relocated.

Source: Intelligent Epistemology §19.1 · §2 · Prop 19.1 · Thm 2.1 · Cor 2.2

Plate 12

Walk each horn and read its ledger, then walk the fourth. Proposition 19.1's claim is exact and narrow: within the consequence structures the MU Theorem covers, the grounding of MU is none of the three. One consistent proposition, one application of reflexivity, one existential introduction (Prop. 19.1, Thm 2.1, Cor 2.2).

The turn

The fourth case

Regress, circle, dogma. Three horns. The trilemma has organised the theory of justification for half a century by insisting those are the options. Proposition 19.1 states the escape precisely: within the consequence structures the MU Theorem covers, the grounding of MU is none of the three. It is a direct existence proof.

The proof has one reflexive inference instance and concludes existentially that at least one consistent inference exists. It terminates. That is the escape. Its premise is an instance of reflexivity, and the chosen proposition disappears under existential introduction. The reflexive analysis of inference begins only after that result is in hand. That order is what keeps the argument non-circular.

The problem of the criterion (Chisholm, 1973) is handled the same way, by role separation rather than by a stronger argument. Method comes from what it is for an answer to come from grounds. Premises, models, representations, and channels supply the material to which the method applies. The circle dissolves and the empirical work of improving the material stays exactly where it was.

What the structure says

Carroll's tortoise, and the rule that will not become a premise

The tortoise accepts a valid argument's premises and demands one further premise before accepting the conclusion. Grant the demand and the same demand reappears immediately, since deriving the conclusion from the enlarged premise set still requires applying a rule. Adding a premise about that application reproduces the role at the next level. The construction iterates forever. No premise ever becomes a rule.

The regress arises from asking a premise to perform the role of the consequence relation. Rules and premises occupy their proper places, and a consequence relation can still be challenged, compared, or revised inside another inferential setting whose grounds and rules are explicit. That closes the regress at the level of rule use.

The other rule-following problem concerns continuation rather than application. A finite record on an infinite domain, an unobserved case, and at least two available values are enough to build two total functions agreeing on everything observed and differing where nothing was. The proposition settles the exact underdetermination present in the record. A unique continuation then needs further structure: a representation, a simplicity or invariance standard, a communal practice, an explicit algorithm.

What the structure says

Experience, and what a signal is worth

Experience is often asked to be an input that supports belief and to carry its own complete interpretation. The first role is available. The second is too strong. The proof of that is a page of arithmetic.

Take an experiential signal and a hypothesis. The evidential force of the signal is the likelihood ratio under the channel and background. Pick two conditional probabilities and the ratio can be made greater than one, less than one, or exactly one. The same sensory token can therefore support the hypothesis, oppose it, or say nothing, and the token alone fixes none of these.

The result locates immediate defeasible support (Pryor, 2000) in the standing relation between signal, background, and channel. An experience enters the problem as a present signal and shifts belief through a standing perceptual channel. Its classification and the channel remain revisable. First-order support and second-order certification are different questions on different levels.

The turn

Indifference returns a shape, not a number

The classical indifference paradoxes are the four shapes of Definition 3.4 in disguise. Take a cube whose side length lies between zero and one. Uniformity in side gives one probability, uniformity in face area gives another, uniformity in volume gives a third. Each is coherent relative to its own description. The three answers reveal a reference family, and the visible constraints preserve all three.

Nor is this special to cubes. Bertrand's chord problem has the same structure, since distinct physical procedures for drawing a random chord induce distinct measures. A physically symmetric die is the contrasting case. Its permutation symmetry among six faces is carried by the problem. The uniform answer is fixed directly. Symmetry carried by the problem itself determines a point. Symmetry that has to be supplied by a choice of description determines a family.

The classification has four outcomes and the paper names all four: a unique reference and hence a point; a family of reference-relative answers; a nonattainment result; or no admissible case at all. The last three become point-valued only when an enriched problem supplies the missing reference, attainment, or consistency structure.

What the structure says

Moore's paradox, and the level the oddity lives on

Moore's paradox lives on the assertion level, not on the proposition. The conjunction of p with the agent's not believing p can be true: it may be raining while an agent fails to believe that it is raining. Suppose a sincere assertion of p is evidence that the speaker believes p. Then a sincere assertion of the conjunction asks the assertion channel to represent the speaker as believing p and as denying that belief at once. The proposition can be logically consistent while its sincere assertion remains unstable as a truthful self-report.

The neighbouring sentence, p but my evidence does not support p, can be coherent when the first clause reports a truth learned through another route. It becomes epistemic akrasia when the speaker treats the second clause as an authoritative assessment of all current grounds and goes on endorsing the first without further reason. The conflict is represented as disagreement between two channels or levels of one model. Treat the higher-order judgement as certain and authoritative, and retaining the conflicting first-order state violates the retention law. Leave it uncertain and the output can remain a joint distribution over the claim and over the reliability of both assessments.

What the structure says

Thresholds, bridges, and the profile

The lottery separates high probability from deductive certainty and from knowledge. A ticket is overwhelmingly likely to lose, and its loss stays graded belief. For every threshold below one there are coherent cases where the probability exceeds it while the background does not entail the claim. The preface runs the other way: an author can rationally be confident in each sentence and expect that at least one is wrong, because conjunction aggregates the small risks its members carry.

Three bridges often conflated with the probability calculus get their own conditional treatments. Dutch-book coherence follows from a fair-price protocol with unrestricted combination and linear valuation. The chance–credence bridge follows from calibration plus admissibility. That pair is the conditional core of the Principal Principle. Reflection follows from the martingale property of a posterior in one model, and it breaks under anticipated forgetting, model change, strategic distortion, or future irrationality.

The section ends by replacing the search for a scalar essence of knowledge with a profile. Truth, internal support, and route integrity are logically independent, so the classical cases occupy different vertices: a lucky guess, a demon world, a Gettier case, a reliable clairvoyant, and ordinary knowledge at the far corner. No analysis based on one coordinate captures the contrasts that motivate the cases, which is the whole argument of a fifty-year literature stated as a cube.

3 horns, and a fourth case
The Münchhausen Trilemma classifies inferential support as regress, circle, or unsupported stopping. MU opens the prior fourth case: a finite direct existence proof that terminates by existential introduction.
premises, and never a rule
Add the premise that A and A → B license B. Deriving B from the enlarged set still uses a rule. A further premise reproduces the role at the next level, and the construction iterates without end.
½ · ¼ · ⅛ three coherent answers
A cube with side length in [0,1]. Uniformity in side gives one half. Uniformity in face area gives a quarter. Uniformity in volume gives an eighth. Each is coherent relative to its description, and the visible constraints preserve all three.
3 coordinates of knowledge
Truth, internal support from the stated grounds, and the reliability or safety of the route across the relevant class of cases. The three vary independently, so no analysis based on one of them captures the classical contrasts.

Prop 19.1 · Thm 2.1 · Cor 2.2 · instrument

The proof terminates. Its premise is an instance of reflexivity from reflexivity and a consistent premise.
Intelligent Epistemology, Proposition 19.1

Verdict Note what is not claimed. MU's proof escapes the trilemma; individual empirical beliefs still need their grounds, their models, and their channels. What the fourth case buys is a floor to stand on while assessing them. The assessment is itself an inference governed by the relation MU concerns. The paper says so explicitly and treats it as a feature.

From the paper
The Münchhausen Trilemma classifies inferential support as regress, circle, or unsupported stopping (Albert, 1968). MU opens the prior fourth case: a finite direct existence proof.

§19.1 · Regress, the Trilemma, and the Problem of the Criterion

Prior is the load-bearing word. The fourth case is not a fourth way of justifying a belief; it is an existence proof that runs before the analysis of justification begins.

The proof idea · Prop 19.1argued

Grounding result

HypothesesThe consequence structures covered by the MU Theorem: a reflexive inferential relation whose language contains at least one consistent proposition.

  1. The proof of MU has one reflexive inference instance, P ⊢ P, for a consistent P.
  2. It concludes existentially that at least one consistent inference exists, so the proof terminates.
  3. Its premise is an instance of reflexivity, taken from reflexivity and a consistent premise, so nothing is assumed unproved at the halt.
  4. The chosen P disappears under existential introduction, so the conclusion does not derive MU from MU.
  5. The reflexive analysis of inference begins only after that result is in hand, which is what keeps the order non-circular.

p. 43 · Prop 19.1 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

p. 52 · §21–§22 — the paper (PDF), opens in a new tab · Induction and knowledge · Part IV 12 / 16

Register — the paper's own margin letters: ·

2 paper sections · §21–§22
  1. §21Induction, Confirmation, and Explanation
  2. §22Knowledge, Testimony, and Scepticism

The trilemma has a fourth case Thirty-one problems, and what follows from each

A right answer can still be bad inference

Induction, the ravens, old evidence, Goodman's riddle, abduction, Gettier, testimony, disagreement, scepticism. Nine famous problems, and one distinction running under all of them: internal support and connection to truth answer different questions. Draw them apart and each problem states its own remaining work rather than resisting solution.

In plain words Being right and reasoning well are two different achievements. You can do one without the other. Most of the classic puzzles about knowledge are built on that gap. Once you stop asking a single notion to cover both, the puzzles turn into ordinary questions about evidence and about how reliable your route to the fact was.

Chapter Seven, Hume’s Ghost — The First Principle Chapter Nine, The New Riddle — The First Principle Chapter Ten, The Best Explanation — The First Principle Chapter Twelve, Do You Really Know? — The First Principle Chapter Thirteen, The Sceptic’s Self-Defeat — The First Principle Chapter Fourteen, Popper’s Wager — The First Principle

ℰ(H) = (T, W, R) — three coordinates, not one essence(0,0,0)(1,0,0) lucky guess(0,1,0) demon world(1,1,0) Gettier(1,0,1) clairvoyant(1,1,1) knowledgeT · truthW · supportR · route integrityNo single coordinate captures the contrasts. Knowledge is the vertex where all three hold at once.

Def 20.2 · Fig. 5 · p. 52

Plate 13

The knowledge profile, after the paper's own plate. Truth, internal support, and route integrity are independent axes, so the classical cases sit at different vertices: lucky guess, demon world, Gettier, reliable clairvoyant, and knowledge at the far corner.

The turn

Induction keeps its standpoint

The classical demand asks for a justification of inductive practice from outside all inference. The reply is short and it is structural. Every reason offered for accepting or rejecting a projective rule is itself an inference governed by a relation of support, so every proposed tribunal is an inference too. Assessment stays reflexively inside the practice. There is no outside.

That settles the constitutive half and hands the rest to empirical work. The future-facing result follows from the projective structure, support, sampling, and distinguishability stated in the problem. Realisability, stability, support, sampling, and distinguishability carry the burden. Where realisability fails, updating can compare only the candidates present in the model class. At best it concentrates on a predictively optimal region under further misspecified-learning conditions.

The paper credits Strawson with the constitutive half of the diagnosis, and then names what his reassurance could not deliver: the convergence conditions. Asking whether induction is rational does resemble asking whether the law is legal. The resemblance does not tell you when a particular projective rule will keep working.

What the structure says

Two paradoxes, one missing model

The ravens paradox arises from combining a logical equivalence with an unstated confirmation model. Evidence confirms a hypothesis exactly when it is more probable under it than under its negation. The likelihood ratio is fixed by the sampling and population model. A nonblack nonraven can therefore confirm, disconfirm, or leave the hypothesis unchanged across different models. Often it confirms minutely, because ravens are rare. The direction is still real.

Sharper still, confirmation follows the hypothesis specified. The paper builds an explicit four-cell model in which evidence confirms a hypothesis while disconfirming a conjunction containing it, which is enough to show that logical containment does not transmit confirmation.

Old evidence is the mirror case. If the agent's current state already assigns probability one to the evidence, conditioning on it does nothing. What is learned when a new theory gains support is the theory-to-evidence relation. The odds move on that relation's likelihood ratio, while the old observation stays inside the current state where it already was.

What the structure says

Representation before projection

The green–grue construction (Goodman, 1955) shows that data support projection only through a represented language. Complexity is representation-relative: a predicate simple in one language may be elaborate in another. A projective inference becomes well defined through its representation language, measurement apparatus, invariances, and candidate transformations. Several permitted representations produce their complete family.

No Free Lunch supplies the computational counterpart in its own finite setting: uniform averaging over an unrestricted problem class equalises the performance of every learner. Successful generalisation therefore requires problem-relative structure. Successful induction makes its bias explicit as exactly that structure.

At deployment scale the riddle recurs as distribution shift. A rule fitted to cases examined before some boundary gets projected onto cases beyond it. The choice among extrapolations lives in representation and apparatus rather than in the record. That sentence is the bridge from a 1955 puzzle to a live engineering problem.

The turn

Gettier, and the third coordinate

Gettier cases show that justified true belief can be true through the wrong connection. The failure is structural. Internal support and connection to truth answer different questions. A case can score on both while the connection between them is accidental. Adding a route condition closes the gap.

A belief-forming route is robust relative to a chosen world class when its kernel keeps discriminating the truth-relevant alternatives throughout that class to the standard the application requires. The choice of class and standard belong to the problem. That is why contextualist and relevant-alternatives insights fit here without a semantic thesis. A route may count as knowledge-producing relative to one class and standard and fail relative to a more demanding pair. The attributions stay consistent because the problems differ.

Two further pressure cases fall out. A subject in a perfectly deceptive demon world can be internally indistinguishable from a normally situated counterpart while lacking the external success. A reliable clairvoyant (BonJour, 1980) with no accessible reason to trust the faculty has the external route without the internal support. Neither coordinate replaces the other. Both are needed.

What the structure says

Testimony, disagreement, and where doubt lands

Bootstrapping cases use a source's outputs to assess that source's reliability, and the channel model names the missing ingredient exactly: an external anchor. Signals identify truth-reliability only through independently verified outcomes or sufficiently restrictive structural assumptions. Repeated self-agreement supplies consistency data. It stops there.

Testimony earns its weight through a Bayes factor in which competence, honesty, dependence, incentives, and selection are separate features. Expert testimony can be rationally weighty without the hearer reproducing the proof, since division of cognitive labour is itself a calibrated channel structure. Dependence matters more than volume. Ten reports copied from one source form one likelihood structure.

Peer disagreement is evidence about another route and often about one's own. An independent, comparably reliable peer supplies a substantial factor. A report determined by the same evidence and method supplies a factor of one and reveals a disagreement to be explained. Conciliation and steadfastness follow from the peer channel, the shared evidence, and the dependence structure in each case rather than from a general policy.

Global scepticism about inference is defeated by the MU Theorem. What remains is local and tractable: Cartesian, brain-in-a-vat, perceptual, testimonial, and scientific doubts become coherent questions about a channel, a model class, or a connection to the external world. A perceptual route may robustly support ordinary claims across ordinary nearby cases while failing to distinguish them from perfectly adversarial simulations. The larger anti-sceptical claim asks more of the route.

4 projectibility conditions
Empirical fit under the channel model. Stability under the chosen transformations. Complexity relative to the chosen coding or reference structure. Predictive performance on held-out or future observations. All four are relative to a fixed representation.
0.145 vs 0.01 the likelihoods that decide it
In the paper's worked model, evidence raises the odds on the hypothesis because 0.145 exceeds 0.01. The same evidence lowers the posterior of the conjunction from 0.45 to about 0.290. Confirmation follows the specified hypothesis rather than the logical equivalence.
1 likelihood ratio, forever
Two model bundles with equal signal laws under every permitted history and action. The product of ratios is one, so no finite history from any adaptive policy moves their prior odds.
4 joint conditions for knowledge
True; internally supported by the problem; route robust in the chosen relevant-alternatives class; the success attributable to the relevant competence or channel. This is offered as a proposal for protection against epistemic luck.
Every proposed tribunal is itself an inference, so inferential assessment remains reflexively inside the practice.
Intelligent Epistemology, §21.1.1

Verdict The pattern repeats with variations. Induction keeps its standpoint and hands the projective task to explicit models. The confirmation paradoxes turn on unstated sampling models. Goodman moves to representation. Gettier gets a third coordinate. Scepticism splits into a self-defeating global version and a tractable local one. In each case what remains is stated rather than dissolved.

From the paper
Gettier cases (Gettier, 1963) show that justified true belief can be true through the wrong connection. A right answer can still be bad inference. The failure is structural: internal support and connection to truth answer different questions.

§22.1 · Gettier, Closure, Context, and Stakes

Three sentences and the middle one is the shortest in the paper. The structural diagnosis is what makes the knowledge profile inevitable rather than stipulated.

p. 60 · §23 — the paper (PDF), opens in a new tab · The map · Part IV 13 / 16

A right answer can still be bad inference Where inquiry gets done

Thirty-one problems, and what follows from each

The section gathers the results in one table. Each row names a classical problem and states what follows, in the paper's own verdict vocabulary: resolved, reconstructed, located internally, reconciled, substantially explained, proved within an experiment class, bounded by the empirical quotient, or separated. Then it classifies the residue into five forms, so what remains has a shape rather than a shrug.

In plain words Here is the scoreboard. Some problems are settled outright. Some are rebuilt so they can be worked on. Some are shown to have a boundary that no amount of thinking will cross. The last part is the useful bit. Whatever is left over falls into five kinds. Each kind tells you what to do next.

A Note on the Companion Paper, A Note on the Companion Paper — The First Principle

The turn

Three patterns under thirty-one rows

Direct resolutions establish inference as its own standpoint, rule application as irreducible to an added premise, thresholds as context-indexed, and channel calibration as externally anchored. Each is a claim that something classical was asking the wrong question. Each comes with the proof that shows why.

Structural relocations move a problem somewhere it can be worked on. Induction goes to projectibility and sampling. Goodman goes to representation. Local scepticism goes to models and channels. Disagreement goes to evidence, dependence, and reliability. Nothing is dissolved by relocation. The problem acquires an address.

Positive reconstructions supply accounts where there was a gap: abduction, empirical knowledge, testimony, scientific inquiry, and the norms internal to inference. These are the rows where the paper builds rather than clears.

What the structure says

Five forms of remaining work

The residue is classified rather than gestured at. Candidate generation: inquiry enlarges the problem with a representation, hypothesis, or rule. Empirical discrimination: new channels or experiments separate the surviving cases. Computational access: further computation reaches more of a fixed full answer. Practical value: value, loss, cost, or an ambiguity rule selects among surviving options. Empirical or metaphysical equivalence: inquiry records the class until a new discriminating structure appears.

Each form has a matching answer type and a matching earlier theorem. A higher-order family. A probability or credal state on an empirical quotient. An approximation contract. An undominated decision set or value-relative completion. A proof of non-identifiability. The classification is useful only because the theorems are already there to receive it.

What the map claims is modest and unusual. It does not say the classical problems have been solved. It says less than that. It says each one has been pushed until it states its own remaining work. The remaining work has five shapes rather than an unbounded number. Five shapes. Generate a candidate. Run a discriminating experiment. Compute further. State the value at issue. Mark the boundary. Each proceeds under the same rule of inference that got the problem here.

Eighteen rows from the map, in the paper's own verdicts (§23.1)
Eighteen rows from the map, in the paper's own verdicts (§23.1)
ProblemWhat follows
Grounding regress and the Münchhausen TrilemmaResolved at the level claimed. MU has a finite direct proof that terminates by existential introduction and precedes the later analysis of inference.
Problem of the criterionResolved by role separation. The inferential form comes from what it is for an answer to follow from grounds; premises, representations, and channels supply the material.
Rule-following and Carroll's regressResolved in two layers. Rule application is irreducible to an added premise, and finite behaviour supports a family of continuations.
The Given and immediate experienceReconstructed. Experience supplies immediate defeasible support through a standing channel and model that connect signal to claim.
A priori support and analyticityLocated internally. Meanings, axioms, and rules determine a priori support within a formal problem; empirical inquiry assesses the framework's relation to the world.
Indifference, reference classes, and self-locationResolved by answer shape. Real symmetry determines a point. Several measures, reference classes, or observation protocols produce their full family.
Uniqueness and permissivismReconciled. One fully defined problem has one full answer, and that answer may contain several permissible credal states.
Dutch books, objective chance, and reflectionResolved conditionally. Fair-price protocols generate Dutch-book coherence; calibrated chance with admissible evidence guides credence; stable Bayesian updating generates reflection.
Moore's paradox and epistemic akrasiaReconstructed through self-representation. The Moorean proposition may be true while its sincere assertion conflicts with the speaker's represented belief.
InductionResolved in two stages. Inference supplies its own standpoint. Projective rules are then compared through explicit models, support, sampling, and distinguishability.
Goodman's new riddle and No Free LunchResolved at the epistemic level. Representation and problem distribution supply the projective structure; evidence compares the cases that make different predictions.
Gettier, closure, context, and stakesResolved through the knowledge profile. Internal support and route integrity are independent coordinates that jointly support empirical knowledge.
Global and local scepticismResolved by scope. MU settles the possibility of inference. Doubt about a channel, model, or world-connection becomes a coherent local problem.
Meno and the swamping problemSubstantially explained. Information has nonnegative expected instrumental value, and knowledge adds a stable, reusable, attributable route to truth beyond lucky arrival.
Duhem–Quine underdeterminationProved within an experiment class. Equal signal laws preserve posterior odds across every available test. New interventions or assumptions enlarge the discriminating class.
Scientific realism, pessimistic meta-induction, and unconceived alternativesBounded by the empirical quotient. Data identify models up to observational equivalence; model enlargement introduces unconceived alternatives.
Logical omniscience and bounded rationalitySeparated. Deductive closure is the full normative answer, while a finite agent reaches a computably accessible subset at each time.
Machine understanding and opacityResolved at the epistemic level and bounded at the phenomenal level. Reliability, route integrity, and auditability determine epistemic standing.

Eighteen of thirty-one, in the paper's order, picked so that every verdict in the vocabulary appears at least once. Ten of the eighteen open with some form of Resolved; in the full table twenty of the thirty-one do. The other seven verdicts are Reconstructed, Located internally, Reconciled, Substantially explained, Proved within an experiment class, Bounded by the empirical quotient, and Separated. The full table runs from page 60 to page 63, and those eight verdicts are eight different claims.

31 classical problems, tabled
From the grounding regress to machine understanding. Each row names the problem and states exactly what follows, in the paper's own verdict vocabulary rather than a summary written for it.
3 broad patterns
Direct resolutions establish inference as its own standpoint, rule application as irreducible, thresholds as context-indexed, and channel calibration as externally anchored. Structural relocations move a problem into representation, sampling, or dependence. Positive reconstructions supply accounts.
5 forms of remaining work
Candidate generation, empirical discrimination, computational access, practical value, and empirical or metaphysical equivalence. The corresponding answers are a higher-order family, a state on an empirical quotient, an approximation contract, an undominated decision set, and a proof of non-identifiability.
The scientific, creative, practical, and metaphysical work remains to be done.
Intelligent Epistemology, §23.2

Verdict The five forms are the paper's own residue classification, not a list of open problems written afterwards. Each has a matching answer type and a matching theorem: candidate generation goes to the higher-order problem, empirical discrimination to the quotient theorem, bounded access to the computational analysis, practical underdetermination to the decision results, and empirical equivalence to the channel model.

From the paper
The gain is knowing what kind of work it is: generate a candidate, run a discriminating experiment, compute further, state the value at issue, or mark the boundary. Each proceeds under the same rule of inference.

§23.2 · Five Forms of Further Work

The closing line of Part IV, and the paper's own answer to the charge that a theory of principle explains nothing. It does not finish the work. It tells you which of five things the work is.

Part V

Science, Computation, and Society

Part V

Six sections of the paper, §24–§29, in one here: falsification, realism and the empirical quotient, bounded rationality, social epistemology, artificial intelligence and causal models, and the scientific cycle.

We can only see a short distance ahead, but we can see plenty there that needs to be done.

Alan Turing, 1950

p. 64 · §24–§29 — the paper (PDF), opens in a new tab · Science and machines · Part V 14 / 16

Register — the paper's own margin letters: ·

6 paper sections · §24–§29
  1. §24Criticism, Falsification, and Confirmation
  2. §25Underdetermination, Realism, and Theory Change
  3. §26Bounded Rationality
  4. §27Social Epistemology
  5. §28Artificial Intelligence and Causal Models
  6. §29A Rational Reconstruction of Scientific Method

Thirty-one problems, and what follows from each The standards arrive with the claim

Where inquiry gets done

Experiments, finite computation, social dependence, and artificial systems. The same distinctions now meet the places where knowledge gets made. Falsification becomes a region of one likelihood comparison. Realism reaches the quotient its experiments fix. Alignment turns out to be reference-based as well as value-relative. And a generative system that reports content its constraints never paid for gets a name.

In plain words This is the part where the machinery meets practice. What a severe test is. How far the evidence lets you go about what is out there. What a bounded reasoner owes when it says it is approximating something. Why ten reports copied from one source are one report. And what exactly is wrong when a model states a fact it never had grounds for.

Chapter One, The Training Run — The First Principle Chapter Fourteen, Popper’s Wager — The First Principle Interlude, The Story of Reasoning — The First Principle Chapter Eighteen, Science Derived — The First Principle Chapter Nineteen, Thinking Together — The First Principle Chapter Twenty, Machines That Reason — The First Principle Chapter Twenty-One, The Transition — The First Principle

The turn

Severity, and what a failed prediction hits

A test is severe for two alternatives to the extent that the resulting signal laws are distinguishable and the apparatus state is controlled. A failed severe prediction can strongly reduce support. A passed one increases support in proportion to how much less expected it was under the rivals. Falsification and confirmation are complementary regions of one likelihood comparison. They are not two epistemologies.

A failed prediction confronts a bundle: theory, auxiliaries, apparatus, and background conditions. Revision localises the change the signal supports and records every alteration of the problem. Popper's insight survives as severe exposure. Duhem's correction is absorbed into a fully specified test package. Kuhn's complication becomes an explicit translation problem, since rival theories may organise salience, measurement, and even the space of candidate questions differently.

What the structure says

How far realism reaches

The observational-equivalence theorem applies directly. Take a class of possibly adaptive experimental policies and call two models equivalent when they induce the same law on complete observation histories under every policy. If both have positive prior probability, every history with positive shared likelihood leaves their posterior odds exactly where they started. Evidence from that class identifies models at most up to the quotient.

The corollary sharpens it. Every property learned from data generated by the class is constant on each equivalence class. Empirically decidable properties are exactly the class-invariant ones. The maximal empirical answer is a probability or credal state on the quotient, together with the structure invariant inside its classes. Structural realism gets a precise statement out of this: where changing theories preserve structure across domains of successful prediction, and that structure is class-invariant, it is the strongest realist content those data support.

Unconceived alternatives get a proof rather than a worry. Here it is. An alternative outside the considered class receives no posterior probability. In an enlarged class, assigning zero prior mass keeps it at zero after every finite update, because Bayes multiplies and renormalises. High posterior concentration on one member therefore establishes comparative success inside the class, and nothing more, unless realisability or class adequacy is separately supported.

What the structure says

What a bounded reasoner owes

An exact theory gives bounded reasoning a target. When a bounded procedure is presented as approximating an exact inferential answer, the claim has to state four things: the exact problem and target answer, the resource budget, an error or regret or calibration or convergence guarantee, and the conditions under which the guarantee holds. A practical procedure may instead be evaluated directly by a loss, a calibration test, or a benchmark. Either way there is a contract. State it.

Deductive closure is the normative answer; the consequences a bounded agent has reached by some time are a subset of it. Further valid computation enlarges the subset without changing the target. That separation makes learning a proof, or a theory-to-evidence relation, genuinely new information for the agent even when it was implicit in the ideal closure.

The paper then applies the contract to a list without exempting anything. A variational approximation states its family, objective, and bound. A Monte Carlo method states its target measure, transition kernel, and mixing control. A heuristic search states its objective, moves, stopping rule, and failure modes. The free-energy principle is treated the same way: a model family with an objective, whose epistemic standing is priced by the channel, reference, and bound it states.

The turn

Dependence, credibility, and what a marker predicts

Agents with a common prior whose posteriors are common knowledge cannot agree to disagree. The scope of that theorem is exactly those two conditions. Absent them, rational disagreement can locate a difference in evidence, reference families, model classes, source dependence, or computational access. Differences in utilities produce disagreement about what to do while the credences agree.

Reports that are conditionally independent given the truth and their source states have a factorising joint likelihood. Reports deriving from one common source, dataset, or coordinated incentive do not. Multiplying them as independent double-counts the evidence. Rational deference therefore depends on competence, independence, conflict of interest, and the hearer's access to calibration data.

The paper states one component of testimonial injustice as a dependence condition. If the truth of a reported claim is independent of a social marker given the source-relevant evidence, then changing a source's weight solely because of that marker makes the answer depend on information with no predictive relevance. Where the marker is predictive only because social structures affect access or treatment, the predictive fact and the moral evaluation of the structure are kept distinct. The formalism represents credibility deficits, testimonial suppression, unequal access, and institutional dependence. Social and moral premises determine their evaluation and remedy.

What the structure says

Machines: prediction, policy, and the standard

Universal induction is indexed to a reference machine. The paper does not soften the point. Once the coding machine is fixed a powerful result follows. Changing the machine changes the case, and adversarial choices can produce pathological universal priors. Leike and Hutter's bad-prior constructions are cited as a decisive control on any finite-scale claim of universality.

AIXI is decomposed into four logically distinct inputs: the algorithmic environment semimeasure, the reward channel, a horizon or discount convention, and expectimax optimisation. Universal induction constrains the first relative to a machine. The other three are separately specified. Prediction and policy stay apart.

Alignment gets the same treatment through the choice theorem. At a fixed decision state the KL-regularised policy is the reference policy reweighted by the exponential of action value, and the components stay distinct: reference support sets the available action class, reward ranks that class, the cost parameter governs departure from the reference, and the optimisation protocol determines implementation. A misspecified or truncated reference can exclude relevant actions. Reward misspecification and reward hacking live in the second slot. Both are old problems in a new place. Too small a cost permits over-optimisation and too large a cost leaves the reference dominant.

Then the standard, applied without exception. A generative system that reports content its constraints never paid for instantiates smuggling at scale. Non-Smuggling applies to machine reporters unchanged. No exemption is available. Hallucination is its violation under thin constraints. Accuracy and auditability remain independent dimensions, so a system may be accurate enough to deserve weight while being too opaque for high-stakes delegation, because action also requires governance, contestability, and responsibility.

Technical · §25 · §28

What this block carriesThe two calculations that carry Part V: the likelihood cancellation behind the empirical quotient, and the bidirectional Gaussian regression.

The quotient cancellation

Empirical equivalence makes the history likelihoods equal under every allowed policy. The policy contributes the same action-selection probability under both models because it is a function of the shared observed history. Bayes leaves the posterior odds at the prior odds. Adaptivity changes the policy without touching the equality.

p. 64 · Thm 25.1 — the paper (PDF), opens in a new tab

Both directions fit

Y = aX + εY, a = Cov(X,Y)/Var(X)

X = bY + εX, b = Cov(X,Y)/Var(Y)

Regress one jointly Gaussian variable on the other and the residual is independent of the regressor. Regress the other way and the same is true. Both structural equations generate the identical joint observational law, so nothing in the distribution orients the arrow.

p. 70 · Thm 28.1 — the paper (PDF), opens in a new tab

The alignment tilt, statewise

π*(a) ∝ πref(a)·exp(Q(a)/τ)

At a fixed decision state the choice theorem gives the KL-regularised policy directly. In a sequential soft-control problem the formula applies statewise once the soft Bellman recursion has determined the action value. The one-step tilt is a component of the sequential construction rather than the whole of it.

p. 69 · §28.4 — the paper (PDF), opens in a new tab

Transcribed from the paper, which marks every one of these proofs optional on a first reading — and prints them anyway. The full text ↗

ℳ / ≡_𝒜 what the evidence identifies
Models are identified at most up to empirical equivalence within the available class of experimental policies. Empirically decidable properties are exactly the class-invariant properties.
0 prior mass, and it stays zero
An alternative outside the considered model class receives no posterior probability. In an enlarged class, zero prior mass survives every finite Bayesian update. That is the force of the problem of unconceived alternatives.
2 directions, one distribution
Every nondegenerate correlated bivariate Gaussian admits linear structural representations in both directions with independent disturbances. Causal direction therefore requires identifying structure beyond the observational distribution.
smuggling at scale the name the paper gives it
A generative system that reports content its constraints never paid for instantiates the failure Non-Smuggling names. The standard applies to machine reporters unchanged. Hallucination is its violation under thin constraints.
A certified boundary is a finding, and a lawful refusal is an answer with its reasons attached.
Intelligent Epistemology, Conclusion

Verdict Two boundaries organise the whole part. Evidence within an experiment class identifies models only up to empirical equivalence, so class-invariant structure is the maximal empirical content. And zero prior mass stays zero under every finite update. A newly conceived alternative gains standing only after the problem is enlarged. Both are proved. Both apply to human and machine inquiry without modification.

From the paper
A generative system that reports content its constraints never paid for instantiates smuggling at scale; the Non-Smuggling standard applies to machine reporters unchanged, and the failure called hallucination is its violation under thin constraints.

§28.3 · Opacity is an audit cost

One sentence connects Corollary 3.5 on page 6 to the most-discussed failure mode in contemporary machine learning. The diagnosis is not that the model is unreliable; it is that the output claims support the constraints never delivered.

What this paper grounds
Within the finite and measurable domains stated in Intelligent Epistemology, dependence on grounds generates probability and exact informational accounting generates relative entropy. A permitted reference and constraints generate least-informative completion; a current state and new constraints generate retention by KL projection.

Intelligent Systems — The Adaptive Closure · Inherited Result 4.1

The descendant states its debt in its own words. Intelligent Systems imports this paper's four laws as an inherited result, adds a value functional and a departure price, and then asks what survives when the resulting operator is composed with itself, with other systems, with observation, with scale, and with time.

Part VI

Norms and the Value of Knowledge

Part VI

If any man is able to convince me that I do not think or act right, I will gladly change; for I seek the truth, by which no man was ever injured.

Marcus Aurelius, Meditations, tr. Long

p. 73 · §30–§33 — the paper (PDF), opens in a new tab · The norms · Part VI 15 / 16

Register — the paper's own margin letters:

4 paper sections · §30–§33
  1. §30Belief, Choice, and Pragmatic Reasons
  2. §31Norms of Inference
  3. §32The Value of Knowledge
  4. §33Internalism and Externalism

Where inquiry gets done The ground asks for nothing

The standards arrive with the claim

Present an output as inference and you are bound, in that capacity, by explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. No moral premise is required to get there. Whether to enter an inquiry at all, how much effort to spend, and how truth ranks among other goods are separate questions with separate premises.

In plain words If you say something follows from your reasons, you have already accepted the rules for that kind of saying — the way calling something a proof accepts the rules of proof. That is not a moral demand imposed from outside. It is what the claim already meant. Whether you should have looked into the question at all is a different matter.

Chapter Eight, The Guillotine — The First Principle Chapter Eleven, The Lottery — The First Principle Chapter Sixteen, The Epistemic Virtues — The First Principle

Three kinds of claim, and which one binds an inference(i)descriptivehow agents in fact reason. An empirical claim about a population.(ii)constitutivewhat must hold for an activity to count as that activity. Present an output asinference and these bind it: explicit grounds, valid dependence, invariance, selection.(iii)practical or morala reason to undertake, continue, or prioritise the activity among competing ends.the standards in (ii) need no premise from (iii)

Def 31.1 · Thm 31.1 · p. 74

Plate 14

Three kinds of claim, kept apart. Descriptive says how agents in fact reason. Constitutive says what must hold for an activity to be that activity. Practical supplies a reason to undertake it. The norms of inference sit in the middle band and need no premise from the third.

The turn

Reasons to act, and reasons to believe

Two agents share a belief state and face different utilities. Their optimal actions may differ while every credence stays the same. A wager, threat, reward, or high stake can therefore change rational action without changing evidential support. The proof takes one line. The hygiene it enforces is worth keeping.

The same distinction clarifies doxastic voluntarism. Agents choose investigations, reports, and actions; the grounds determine support. Ordinary use of the word knowledge may stay sensitive to pragmatic context. The line between support and stakes does not move with it.

Decision paradoxes are handled the same way. Newcomb's problem requires the relation among prediction, causation, policy, and utility to be stated, and evidential and causal decision theories formalise different dependence structures. The supplied structure determines which problem is being solved. That is not a dodge. It is the diagnosis.

What the structure says

Constitutive, not moral

Theorem 31.1 makes the claim precise. An agent or process that presents an output as inference is bound, in that capacity, by the standards of Part I: explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. An output also presented as complete must preserve the whole supported answer.

The proof runs through the method theorem. An output satisfies the inferential description exactly when every dependence it carries is carried by its stated grounds. The standards are therefore internal to the claims being made, in the way the rules of proof bind a purported proof and preservation laws bind a purported homomorphism.

A reasoned rejection of every inferential norm is itself offered as an inference. The epistemic standpoint is inescapable for reasoned evaluation. What that inescapability does not settle is whether to enter an inquiry, how much effort to spend, or how epistemic goods trade against non-epistemic ones. Those need an independent premise. The paper says so plainly rather than reaching for one. It does not overclaim at the end.

What the structure says

What knowledge is worth

Free optional information never hurts. After observing any signal, the best action available is at least as good as retaining the action already chosen. Taking expectations gives the inequality directly. With a cost attached, inquiry is strictly preferred when the value of the information exceeds it, rejected when it falls short, and indifferent at equality. The result is exact for the supplied utility and cost.

The Meno problem asks why knowledge is worth more than mere true belief, since both can guide the same immediate action. The route-integrity answer is that knowledge is stable under perturbation, reusable across nearby problems, and attributable to a method. A true answer reached accidentally may guide one action. Once. A truth-connected method supports navigation, transfer, and correction.

The swamping problem asks what a reliable route adds once truth is already present. Truth is the successful outcome; a robust route explains why the success recurs, survives relevant counterfactual changes, and can be credited to the agent or the channel. That supplies the extra epistemic value. Final moral value stays in the practical domain where it belongs.

Internalism and externalism close the part by answering different questions. It is the pattern of the whole paper compressed into two pages. Internal support asks whether the state is a valid answer to the agent's grounds. External success asks whether representation and channels connect those constraints to the world in the required way. A demon-world subject may be internally impeccable and externally unsuccessful; a lucky guesser may be externally correct and internally unsupported. Reliabilist, safety, competence, and virtue accounts supplement the internal analysis rather than replacing it.

3 kinds of claim
A descriptive claim says how agents in fact reason. A constitutive claim states what must be true for an activity to count as that activity. A practical or moral norm supplies a reason to undertake or prioritise it.
4 standards that come with the claim
Explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. An output also presented as complete must preserve the whole supported answer.
V₁ ≥ V₀ free optional information
After any signal the inner maximum is at least the conditional expected utility of the action already chosen. Taking expectations gives the inequality. Information that is free and optional has nonnegative optimal expected utility.
V₁ − V₀ > c when inquiry is required
Strictly preferred above the cost, rejected below it, indifferent at equality. The result is exact for the supplied utility, cost, and option to ignore the signal. Categorical moral duties belong to the wider value problem.
A reasoned rejection of every inferential norm is itself offered as an inference and is therefore assessable by inferential standards.
Intelligent Epistemology, §31

Verdict The epistemic is–ought is resolved constitutively rather than derived. The standards of inference bind outputs offered as inference, in the way the rules of proof bind a purported proof and preservation laws bind a purported homomorphism. An independent normative premise is needed only for whether to enter the inquiry, how hard to work at it, and how epistemic goods trade against others.

From the paper
An agent or process that presents an output as inference is bound, in that capacity, by the standards of Part I: explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem.

Thm 31.1 · Norms of inference

In that capacity is the whole qualification. The theorem binds the output, not the agent's life. It says what a claim of inference has already committed to. No premise about what anyone ought to care about is needed.

The proof idea · Thm 31.1argued

Norms of inference

HypothesesAn agent or process that presents an output as inference.

  1. The method theorem analyses what it is for an answer to count as following from its grounds.
  2. An output satisfies that description exactly when every dependence it carries is carried by its stated grounds.
  3. An output also presented as complete must additionally preserve every supported case and every supported distinction.
  4. The standards are therefore internal to the claim being made, not added to it from outside.
  5. A reasoned rejection of every inferential norm is itself offered as an inference and falls under the same standards.

p. 74 · Thm 31.1 — the paper (PDF), opens in a new tabcompressed from the paper’s own optional proof

The sibling route
Within the agent domain and consistency requirements stated in Intelligent Economics, bounded valued comparison against a reference forces the same exponential choice law, its score decomposition, its KL-regularised variational dual, and a canonical reversible relaxation.

Intelligent Systems — The Adaptive Closure · Inherited Result 4.2

Intelligent Economics reaches the same choice law from the other side, starting from bounded agents comparing options against inherited expectations rather than from what an answer may contain. Two derivations, different premises, one operator. The keystone paper imports both and asks what survives composition.

§23.2 · Five forms of further work · Part IV

What is left, and what kind of work it is.

The paper does not claim the classical problems are finished. It claims each has been pushed until it states its own remaining work, and that the remaining work has five shapes. Each shape below carries its answer type and the theorem that receives it.

  1. (i)

    Candidate generation

    Inquiry enlarges the problem with a representation, hypothesis, or rule. Nothing inside a fixed candidate class can reach a candidate outside it.

    The answer is a higher-order family.

  2. (ii)

    Empirical discrimination

    New channels or experiments separate the surviving cases. Until one arrives, the equivalence class is the result.

    The answer is a probability or credal state on an empirical quotient.

  3. (iii)

    Computational access

    Further computation reaches more of a full answer that was already fixed. Deductive closure is the normative object; the accessible subset is where a finite agent stands.

    The answer is an approximation contract.

  4. (iv)

    Practical value

    Value, loss, cost, or an ambiguity rule selects among surviving belief or decision options. Evidence does not supply them.

    The answer is an undominated decision set or value-relative completion.

  5. (v)

    Empirical or metaphysical equivalence

    Inquiry records the equivalence class until a new discriminating structure appears. The record is itself a finding.

    The answer is a proof of non-identifiability.

p. 77 · Conclusion — the paper (PDF), opens in a new tab · The ground · Conclusion 16 / 16

The standards arrive with the claim The paper, from its conclusion back to §1

The ground asks for nothing

One fact at the base: consistent inference is possible. Everything above it is what that fact contains once each domain has named its objects. The principle adds nothing to any problem it meets. By adding nothing it lets each problem show exactly what it holds.

In plain words The whole book rests on something so small it is almost embarrassing to state: reasoning can be done consistently. That is it. Every law, every resolution, every boundary in these pages is what follows from taking that one fact seriously inside a properly stated question.

Epilogue, One Foundation — The First Principle Coda, The Return — The First Principle

MU is the floor of this account. The floor is thin by design. Consistent inference is possible; nothing more is claimed at the base. Everything built above it is what that one fact, taken seriously inside each stated domain, turns out to contain. An answer offered as following from its grounds must draw on those grounds alone. Every answer-relevant dependence therefore lives in the problem. Equivalent descriptions agree. A complete report keeps the whole supported result.

Held to arithmetic, the discipline takes four forms. Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Under finite counting symmetry it is Shannon's maximum entropy. Revision is KL projection, carrying forward every feature the new constraints permit. Values and costs then govern choice, channels carry belief's connection to the world, and four world-facing conditions decide when that connection finds the truth. These are laws rather than customs. Within their stated domains, no consistent alternative exists.

What the structure says

What the problems became

Regress ends at the exhibited floor. Indifference returns a point, a family, or a certified boundary according to the structure present. Induction becomes a set of explicit projective and world-facing conditions with the price of each on the table. Goodman moves projectibility into representation. Gettier joins internal support to route integrity. Scepticism separates the self-defeating universal doubt from the tractable local kind.

Scientific realism reaches the quotient its experiments fix, and the limit theorems contribute their ceilings and empty classes as complete answers in their own right. A certified boundary is a finding. A lawful refusal is an answer with its reasons attached. Both are results.

What the structure says

What is left

Work with a shape, and only five of them. Generate a candidate, design a discriminating experiment, compute further, state the value at issue, or record the equivalence that survives. Each task is the same discipline carried into new ground. Nothing new is needed.

And the ground itself asks for nothing. It was there before the argument began, and every argument, including any brought against this one, stands on it.

Π ⟼ Answer(Π) the whole operator
The problem gives the subject matter. MU returns its complete answer. The structure present fixes the resolution: a point, an orbit, a wider family, or an empty requested class certified by proof.
4 laws, none with an alternative
Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Revision is KL projection. Within their stated domains, no consistent alternative exists.
1 assumption at the base
Consistent inference is possible. Nothing more is claimed at the floor. The thinness is the point: every substantive condition stays visible above it.
One rule for what follows. Four quantitative laws. Every answer shape.
Intelligent Epistemology, the paper's closing line

Verdict Held to arithmetic, the discipline takes four forms and each is unique in its domain. Met by the classical problems, it resolves each at exactly the level its grounds support. What remains has five shapes. And the floor was never in question: every argument, including any brought against this one, stands on it.

From the paper
It was there before the argument began, and every argument, including any brought against this one, stands on it.

Conclusion · p. 77

The paper's last sentence before its closing display. Read it against Corollary 2.2 on page 4. There the denial supplies the witness. Eighty pages later the same move closes the book.

In one sentence The conclusion, p. 77 — the paper (PDF), opens in a new tab

It was there before the argument began, and every argument, including any brought against this one, stands on it.

Afterword · The descendants

What this paper grounds.

Intelligent Epistemology settles what an answer may contain when it must follow from stated grounds. Nothing below MU is assumed. Intelligent Economics reaches the same operator from the other end: bounded systems choosing under inherited expectations, priced by the cost of departing from what they already carry, arriving at the identical exponential form by a route that never mentions grounds. Intelligent Systems imports two of the results below verbatim and asks what survives composition, observation, scale, and time. The First Principle carries the discipline in book voice.

Each object this paper settles, and the work that carries it
Object Settled here Carried into
MU the floor consistent inference exists, proved in one line from reflexivity, before any of the analysis that uses it The family the base every later paper stands on. Nothing in the family is assumed below it. No argument against it can be made without it.
q dependence on grounds an answer rule depends only on the problem exactly when it factors through the quotient, F = F̂ ∘ q Systems the invariance its closure theorem is stated under: a result that changes with a relabelling was never a result about the system.
D relative entropy exact marginal–conditional decomposition forces DKL up to scale, and prices the epistemic state the constraints support Systems the observation channel, carried across verbatim rather than reproved, then read as an effective drive under coarse-graining.
A full answers a complete report keeps every supported answer and every supported distinction: a point, an orbit, a family, or a certified empty class Systems the discipline of leaving a family unresolved, and the record that keeps a committed action revisable.
μ the permitted reference least-assuming completion is relative-entropy projection from a permitted reference, with Shannon maximum entropy as its finite counting-symmetry form Economics the inherited expectation a bounded system already carries. The sibling route arrives at the same object without ever mentioning grounds.
τ value and its price with belief fixed, values and costs govern choice: the tilt eV/τ is exactly what that pair forces Economics the economic law itself, with τ read as the shadow price of information rather than as a free parameter.
0 Epistemic Zero add only supported structure, and preserve every supported distinction — the bilateral rule the whole method reduces to The First Principle the same discipline in book voice, without the apparatus. Every section of this edition names the chapters that tell its part of it.
G generation a higher-order problem can hold a family of candidate languages, representations, and hypotheses; inventing a member of that family is not an inference Not carried the one object this paper does not supply. Intelligent Systems makes generation a separate operator, priced by selection after the fact.

Nothing in the right-hand column is reproved downstream. Each work names its import, cites the result, and carries it across. A claim that cannot name where its structure entered is the claim this paper calls smuggled. That is the whole test.