MU is the paper's name for one principle: consistent inference is
possible. It is the thinnest fact reasoning can stand on. One line of
reflexivity proves it. Everything above that floor turns on the word
from: a conclusion follows from its grounds exactly insofar as
every answer-relevant dependence is carried by those grounds. Labels,
coordinates, search order and priors earn inferential force in one way
only. The problem names them. Assume nothing the constraints omit.
Surrender nothing they deliver.
Interactive research edition · six parts, thirty-three sections · every
claim carries its § or theorem number
Contents
The argument
Thirty-three sections, folded into sixteen.
Nothing is dropped. Each row names the paper part, the sections this edition makes of it, and how many numbered statements and argument blocks sit underneath. From a one-line existence proof to the norms of inquiry.
When you know a thing, to hold that you know it; and when you do not know a thing, to allow that you do not know it: this is knowledge.
Confucius, Analects 2.17, tr. Legge
Before the argument
One fact at the base, and everything it turns out to contain.
MU begins with the thinnest fact on which reasoning can stand: consistent inference is possible. The proof takes one line, since for any consistent proposition P, reflexivity gives P ⊢ P. Everything else turns on the relation the
word from names: a conclusion follows from its grounds
exactly insofar as every answer-relevant dependence is carried by
those grounds.
Intelligent Epistemology, abstract
MU ≡ ∃I CI(I)
F = F̂ ∘ q
MU : Π ⟼ Answer(Π)
The paper boxes all three. The first is the floor, the second is
what any inference commits itself to, the third is the operator
the first two produce.
The paper runs 33 numbered sections in
6 parts over 81 pages, closing on an
unnumbered conclusion. This edition maps them into 16 sections
in six, and every section prints the paper range it covers in its eyebrow.
Underneath sit 157 numbered statements and remarks,
together with 49 argument blocks. Not one argument
block appears before §19.
How claims are graded on this site
Results are graded by their statement form throughout: theorems, propositions, and corollaries are proved; arguments are defended on stated grounds without claim of proof; remarks locate, delimit, and connect.
Intelligent Epistemology, §1.1
provedTheorem, proposition, corollary, or lemma. Proved.
arguedAn argument block. Defended on stated grounds, without claim of proof.
locatedA remark. It locates, delimits, and connects.
definedA definition. It fixes a term and asserts nothing.
openThe paper marks this one unsettled.
That sentence is the whole of the grading, and this site copies it
rather than inventing a scale. A claim carries the register of the
statement form it came from and no other. An argument summarised
here stays argued however persuasive the argument is. Every
tagged figure names the section, theorem, proposition, corollary,
lemma, definition, or remark it comes from, and every number links
to the page that states it.
What MU returns, object by object
∃I
consistent inferenceDef 1.3↗The floor. At least one valid passage from a consistent premise set exists, and the identity instance exhibits it.
Π
the inference problemDef 3.1↗Seven slots: representation language, candidate domain, premises and constraints, consequence relation, question, required form of answer, and the rule for when two outputs count as the same.
A
the full answerDef 3.3↗Everything the problem supports, together with the equivalences it declares. A point, an orbit, a family, or a certified empty class.
E
the inference episodeThm 4.1↗Rule, content, and operator. Remove any one and no episode remains, though the abstract relation survives without the third.
∅
the certificateProp 6.1↗An empty admissible class is a complete classification when the proof of emptiness comes with it. Gödel, Turing, Tarski, Fitch, and Arrow all return one.
p
probabilityThm 9.2↗MU's unique form for determinate scalar plausibility on a Boolean event structure, up to order-preserving regraduation.
D
relative entropyThm 10.5↗The unique measure of total informational change whose staged and flat descriptions balance, up to positive scale.
μ
the two projectionsRemark 10.2↗One divergence, two reference objects. Completion projects from a permitted reference; revision projects from the current state.
𝒦
the credal familyDef 11.4↗The set of point-valued answers a problem permits. Where it holds several members, the set itself is the answer.
K
the evidence channelDef 16.1↗The stochastic kernel that carries the connection to the world. A single reliability number is only a special summary of it.
This paper prints no symbol table. It introduces its vocabulary
where it needs it, in definitions, and the list above is that
vocabulary in the order the argument builds it. The rail beside you
tracks the same ten objects, lighting each as the section that
declares it arrives.
These minutes are counted off the page rather than guessed. The skim
layer — headline, dek, plain box and verdict, sixteen times — runs
about 10 minutes. The argument itself runs about
60. Opening the 6 technical
blocks and the 8 proof ideas adds roughly
10 more. There is no
layer control on the page. These are depths of reading rather than
modes to select, and the paper prints no reading guide to build one
from. The contents plate above is the shortest route in: six part
rows, each naming the paper range it folds.
1line of proof
For any consistent proposition P, reflexivity gives P ⊢ P. The premise set is consistent, the instance is valid, and existential introduction discharges the chosen P. Everything in the paper stands above that one line.
Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Revision is KL projection. Within their stated domains, no consistent alternative exists.
Every argument carries a promise: the conclusion comes from its grounds. The paper takes that promise at face value. If an answer is presented as following from stated grounds, then any feature that changes the answer must appear among those grounds or stay visible as unresolved freedom. Beneath the demand sits a floor thin enough to prove in a line. Consistent inference is possible.
In plain words When you say a conclusion follows from your reasons, you are making a claim you can be held to. Anything that changes your answer has to be one of the reasons you gave. If it is not, you were carrying something you never declared. The whole book is that one demand, taken seriously.
There are two ways to write a theory of reasoning. A constructive account builds inference out of proposed mechanisms and asks whether the machine works. A theory of principle starts from a constraint every genuine inference must satisfy and asks which forms meet it. This paper takes the second road. The road has a shape: state the problem, determine whether an answer exists, classify every answer up to the relevant equivalence, and require that a mere redescription leave the result unchanged.
The paper grades itself as it goes. Theorems, propositions, and corollaries are proved. Arguments are defended on stated grounds with no claim of proof. Remarks locate, delimit, and connect. That sentence in §1.1 is where this site gets its register marks. No claim here is graded higher than the statement form it came from.
What the structure says
The word from
Consider what the ordinary word is doing. An answer presented as following from its grounds has made a commitment about dependence. The commitment is checkable. Take any feature that would change the answer. It appears among the grounds, or the answer had help.
Labels, coordinates, orderings, priors, and preferences are the usual stowaways. Each is harmless once written down and given a role. None is harmless while it sits outside the stated problem and steers the result anyway. The commitment attaches to the relation rather than to any particular vocabulary for it, which is why the argument survives translation into other formal settings.
What the structure says
What MU says, and what it does not
MU is the principle that consistent inference is possible. In symbols it says only that there exists an I such that I is a valid inference from a consistent premise set. It does not say what follows from what. It does not privilege a logic, a language, or a method. It claims existence. Nothing more.
That thinness is deliberate. A floor thick enough to be interesting would be thick enough to be denied. The paper wants a base nothing built above it can undermine. Definition 1.1 keeps the relation apart from the episode: the formal core concerns the inferential relation, while claims about agents concern actual uses of it. The distinction pays off two sections later.
What the structure says
The proof, and the shape of the denial
Let P be any consistent proposition. Reflexivity gives the valid instance P ⊢ P. Its premise set is a single consistent proposition, so the instance is a consistent inference, and existential introduction discharges the chosen P from the conclusion. The theorem is about the kind rather than the witness. That is the whole proof.
The corollary runs the same machinery backwards. Suppose a language contains the proposition that no consistent inference exists. Suppose further that the proposition were consistent. Reflexivity would immediately make the identity inference from it to itself a consistent inference. The denial exhibits what it denies. This is stability under self-reference, not a rhetorical trap: the proof of the theorem never needed the denial, and the corollary only records what happens when someone tries.
Two movements of one principle follow: existence, then form. The MU Theorem establishes that a consistent inference exists. The Dependence Theorem, three pages later, unfolds what any such passage commits itself to once its conclusion is said to come from its grounds. Everything quantitative in the paper is downstream of that second movement. Existence first. Form after.
1line of proof
For any consistent proposition P, reflexivity gives P ⊢ P. The premise set is consistent, the instance is valid, and existential introduction discharges the chosen P. That is the whole proof of the floor.
Suppose ¬MU were consistent. Reflexivity would then make the identity inference from ¬MU to ¬MU a consistent inference, witnessing the existence the proposition denies.
Verdict The floor is not a postulate. Theorem 2.1 exhibits a consistent inference rather than assuming one, and Corollary 2.2 shows the denial cannot be consistently asserted, because reflexivity turns the denial itself into the witness it denies. Everything after this is what the demand contains once a domain names its objects.
▸From the paper
An answer that changes with an unstated label, coordinate, ordering, prior, or preference draws on more than the grounds supplied.
The whole discipline in one sentence, stated before any machinery arrives. Every later theorem is this demand met inside a domain that has named its objects.
▸The proof idea · Thm 2.1argued
Existence of consistent inference
HypothesesA reflexive inferential relation whose language contains at least one consistent proposition.
Let P be a consistent proposition. The hypothesis says there is one.
Reflexivity gives the valid instance P ⊢ P.
Its premise set is {P}, which is consistent by choice of P.
So P ⊢ P is a consistent inference, and existential introduction discharges P.
The result is about the kind, not the witness: the chosen P disappears from the conclusion.
Assume nothing beyond what the constraints demand.
The First Principle · Reasoning (in) the Intelligence Age
The trade book runs the same argument for a reader with no formal training, opening on the same Confucius line this paper's Part I carries. It grades its claims in three bands rather than five: proved, best explanation, avowed.
provedThe proof, at full lengthOne consistent proposition, one application of reflexivity, one existential introduction. Then the denial, which reflexivity turns into the witness it denies.≈ 20s
The proof at full length, and the denial that supplies its own witness. One consistent proposition, one application of reflexivity, one existential introduction (Thm 2.1, Cor 2.2). The fourth panel walks the Münchhausen horns and shows why none of them is this (Prop 19.1).
The answer depends on the problem and on nothing else
Write the working presentation of a problem and the problem itself as two different things, joined by a map that forgets labels, coordinates, enumeration order, and temporary gauges. An answer rule depends only on the problem exactly when it factors through that map. Theorem 3.1 turns an ordinary word into an equation, and the equation generates the rest of the paper.
In plain words Two people can write the same problem down in different ways. If their answers differ, something in how they wrote it is doing work. That something was never declared. The fix is to say what counts as the same problem, then require the answer to depend only on that. Once you do, most of the argument writes itself.
A complete answer is whatever the problem supports, kept whole. There
are four shapes it can take. This card opens on the fourth: an optimum the case
approaches and never reaches. Move the slider and the divergence falls
toward an infimum no member attains. The other chips declare a
symmetry, supply a second reference, or tighten a constraint until the
answer changes shape.
Π ⟼ Answer(Π)point · orbit · family · certificate
the open constraint · shape (iv) · D = 0.130812 nats at p₁ = 0.7500 · infimum 0, attained by nobody
the problem declares
Definition 3.4, four shapes. This problem returns:
(i)one class, one representative, fixed by every symmetry
(ii)one orbit, its representatives exchanged by the symmetry
(iii)more than one inequivalent answer, retained as the family
(iv)an empty class or unattained optimum, with its certificate
members∞inequivalent answers retained
H(p)0.5623nats, at the position shown
DKL(p‖μ)0.1308departure from the uniform reference
boundary(½, ½)the point no member of the class reaches
The infimum of D(p‖μ) over the open case is 0, approached as p₁ falls to ½
and attained by no member, because the minimising point lies on the
excluded boundary. At p₁ = 0.7500 the divergence reads 0.130812 nats, computed from the closed form and again from the sum
over both atoms.
The two disagree by exactly zero. Push the slider down and the number falls
without ever arriving. The complete answer is three things: the case, the
infimum 0, and the boundary point (½, ½) that no member reaches.
Shape is not a matter of how hard the problem is. It is a matter of what
the problem carries. Every move that turns a family into a point on this
card is a declaration. The card names it.
defined the four shapes are
Definition 3.4 and the orbit rule is Corollary 3.3; the four presets are
the paper's own worked examples in §11.4, recomputed here rather than
quoted.
argued the translation case is
drawn on a finite window of at most 200 cells. That clamp is the
frame's, not the paper's: on ℝ the invariant class is Lebesgue, the total
mass is infinite at every finite window and beyond it, and no normalised
invariant probability exists at all.
▸The measured residual
closed form against the atom sum
0
Each of these is one quantity computed two ways from the state on
screen, with neither route derived from the other. A figure of 0 means
the two routes landed on the same double. Anything else is the width of
one. Where a line above reports how far two readings disagree, this is
the figure it stands for.
Move the constraints and the symmetry, and watch the answer take one of Definition 3.4's four shapes. The four settings on the right are the paper's own worked examples from §11.4: a symmetric die, two standing references, translation on the line, and an open constraint whose optimum is never attained.
The turn
Writing the problem down
A problem is more than its data. It is a specification. It includes the language the grounds are expressed in, the candidates on offer, the constraints, the rule of consequence, the question being asked, the form of reply required, and the rule for when two outputs count as the same answer. Definition 3.1 makes all seven explicit and calls the whole specification an inference problem.
Any one of the seven may be stipulated, inherited from a domain, or supported by a further argument. Writing it down identifies its role and exposes whatever supports it. A language, consequence relation, reference measure, output form, or equivalence offered as a conclusion becomes the answer to a higher-order problem. It follows the same rule.
What the structure says
The equation
Let 𝒳 be the class of working presentations and 𝒫 the class of problems, joined by a surjection q that forgets everything the problem treats as irrelevant. An answer rule assigns to each presentation an answer in the answer space of its problem. Theorem 3.1 then says something short. The rule depends only on the problem exactly when it is constant on every class q identifies. That is exactly when it factors as F = F̂ ∘ q.
One direction is immediate. The other constructs F̂ by picking any presentation of a given problem and reading off its answer, then uses constancy to show the choice does not matter. A second condition rides alongside. Across a structure-preserving redescription, answers must transport with it. Mathematicians call that naturality.
Put two reasoners in front of the same problem. A difference in their answer spaces reveals a difference in grounds, representation, or procedure. Naming the difference puts it in the problem. Until it is named, the full answer keeps every possibility the shared grounds support. Nothing is discarded on trust.
What the structure says
What completeness costs
Definition 3.3 fixes the other half. A report is complete when it preserves the entire answer space, together with the equivalences the problem declares. Completeness is exhaustion relative to the question. Computability, decidability, and single-valuedness are separate properties. None is required.
Theorem 3.2 then closes the loop. A result presented as complete preserves every satisfying answer, every distinction the equivalence leaves inequivalent, and every unresolved family. A question, output form, equivalence, or selector compresses the result precisely when it makes the corresponding distinction irrelevant. Those data define the richer problem whose complete answer is the compressed one. Nothing is lost by insisting on this. Something is gained: the compression becomes visible as a further piece of structure rather than a silent convenience.
What the structure says
Canonical points and the symmetric pair
A point is canonical exactly when the problem contains structure that distinguishes and fixes it. Suppose a symmetry of the problem exchanges two satisfying answers. Every point rule that follows from the problem is invariant under that symmetry, so the invariant output retains the exchanged answers together. The full answer is the pair. Not one of them.
Enlarging the problem by an order supplies the distinction that selects one of them. The order is then part of the record. Its own standing depends on its own grounds. This is a small result with a long reach: it is why the paper never has to choose between a point and a family. The grounds choose.
The turn
Epistemic Zero, and the axiomatic method as a special case
The rule is bilateral. Two halves. Add only supported structure. Preserve every supported distinction. In plain words: assume nothing beyond what the constraints demand, and surrender nothing they deliver. Drop the first half and unstated assumptions creep in; drop the second and answers get compressed into points nobody paid for.
The ordinary axiomatic method falls out as an instance. Specify a language, an axiom theory, a consequence relation, a question, a form of answer, and an equivalence, and the complete theorem set is deductive closure. Nothing exotic happens here. The general rule becomes the familiar one because the familiar one already satisfies it.
The four shapes a complete result can take (Def. 3.4)
The four shapes a complete result can take (Def. 3.4)
Shape
What the problem has done
What the report must carry
A point
Its structure distinguishes and fixes one representative under every symmetry
the representative, and the structure that fixed it
An orbit
A symmetry fixes the answer-class while leaving its representatives equivalent
the orbit, not one member of it
A family
More than one inequivalent answer survives the stated grounds
the whole family, as the answer
A certificate
The admissible class is empty, or the optimum is not attained
the case, the boundary, and the proof
The first two shapes separate uniqueness of the answer-class from whether the problem singles out a representative. A selector supplied later does not correct the earlier answer; it defines a richer problem whose complete answer is the compressed one (Thm 3.2).
F = F̂ ∘ qdependence on grounds
An answer rule depends only on the problem precisely when it is constant on every class of presentations that q identifies. Labels, coordinates, enumeration order, and temporary gauges are exactly what q forgets.
Π = (ℒ, ℋ, 𝒞, ⊢, 𝒬, 𝒪, ∼): representation language, candidate domain, premises and constraints, consequence relation, question, required form of answer, and the rule for when two outputs count as the same answer.
One equivalence class with a representative fixed by every symmetry; one orbit whose symmetry fixes the class while leaving its representatives equivalent; more than one inequivalent answer, retained as the family; an empty admissible class or unattained optimum, with its certificate.
Non-Smuggling stated positively. Content fidelity: each proposition tracks the support its premises carry. Selection fidelity: each point choice tracks a selector the problem carries. Calculus fidelity: each operation tracks the structure the problem supplies.
The structure present fixes the resolution of the result: a point, an orbit, a wider family, or an empty requested class certified by proof.
Intelligent Epistemology, §3
Verdict The factorisation is the whole engine. Everything from probability to the treatment of Gettier cases is this one equation applied where a domain has named its objects. Completeness comes with it: a report presented as complete preserves every satisfying answer, every distinction the equivalence leaves standing, and every unresolved family. Compression is legitimate only when the problem itself makes the compressed distinction irrelevant.
▸From the paper
An answer follows from a problem exactly when every distinction that changes the answer survives in the problem.
The paper sets this line off on its own between the theorem and its discussion. It is the theorem in English. The rest of the argument keeps returning to it.
▸The proof idea · Thm 3.1argued
Dependence on grounds
HypothesesA surjection q from working presentations onto problems, and an answer rule F assigning to each presentation an answer in the answer space of its problem.
Suppose F factors as F̂ ∘ q. Then two presentations of one problem have the same image under q and therefore the same answer.
Conversely, suppose F is constant on every class q identifies. Define F̂(Π) as F(x) for any presentation x with q(x) = Π.
Constancy on the class makes that definition independent of which x is chosen.
Surjectivity of q gives existence for every problem, and the construction gives uniqueness.
The transport equation extends the same demand from alternative presentations to equivalent descriptions of the problem itself.
An abstract inferential relation needs rules and content. An actual episode also needs something that carries the relation out. Keeping the three apart stops a proof being confused with the act of proving it, and it lets the foundation theorem state the order of business without circularity: MU is proved first, and the method is applied afterwards, including to MU.
In plain words There is a difference between a rule of logic, the thing being reasoned about, and whoever is doing the reasoning. Mix them up and you get puzzles that are not puzzles at all. Keep them apart and an old problem — do you need a method before you can pick out good cases, or good cases before you can pick out a method? — stops being a circle.
Logic, Content, Agent. The three failures around the outside are the proof: rule and content without an operator leave an abstract relation, rule and operator without content leave execution with nothing inferred, content and operator without a rule leave a change of state.
The turn
Logic, Content, Agent
Logic is the rule structure that separates valid from invalid transitions and lets inferences compose. The domain under study fixes which consequence relation is in play. Content is the subject matter represented in premises and conclusion, together with whatever makes the former bear on the latter. Symbol manipulation can instantiate a calculus; reading it as inference about a domain requires a semantics.
The agent or operator is whatever carries the transition out in an actual episode. A person. An institution. A proof assistant. An algorithm. Anything that does it. The theorem concerns that function and says nothing about consciousness. The result stays available for machine reasoners without smuggling in a philosophy of mind.
What the structure says
Why the three cannot be pulled apart
Theorem 4.1 states an equivalence: an episode is an inference episode exactly when all three aspects are present. The forward direction reads them off the description itself, since from supplies the rule, premises and conclusions supply the content, and drawing supplies the operator.
The converse is where the work is. It proceeds by watching each separation fail in a different way. Take away the operator and an abstract relation remains, with no event. Take away the content and rule execution remains, with nothing inferred. Take away the rule and a change of state remains. A change of state is not a passage from premises to a conclusion.
Remark 4.1 keeps the scope honest. The theorem is relative to the object being analysed. An abstract inferential relation contains rules and content and needs no operator; an episode adds one. Definition 1.1 keeps consequence relations available at the relational level throughout.
What the structure says
The foundation theorem
Three clauses. The order among them is the content. Proof order: MU is proved independently of the later method, which is then applied to every subsequent proof, question, and assessment. Minimality: any proposition sufficient to ground the existence of consistent inference entails MU, so MU is the weakest claim sufficient for the task. Self-application: MU has a direct proof, while every valid assessment of it is itself an inference governed by the relation MU concerns.
Reflexive application follows the proof rather than replacing it. That sequence is what distinguishes this from a bootstrap. A theory that proved its floor by using its own method would be arguing in a circle; a theory that proves its floor first and then applies the method above it is doing ordinary mathematics.
What the structure says
Criterion and content
The problem of the criterion asks whether one must begin with reliable cases and infer a criterion, or begin with a criterion and identify reliable cases. Stated that way it looks circular. It is not. It stops looking circular once method and subject matter occupy their proper roles.
The method follows from what it is for an answer to come from its inputs: relevant information is visible, equivalent descriptions agree, and a complete report preserves the full supported answer. Premises, models, and channels provide the material being assessed. Experience tests their fit to the world. The inferential form governs the assessment, and it does not compete with the material for the same job.
Fallibilism survives intact. A criterion can govern what follows from a model while the model, the representation, or the channel stays revisable. The circle dissolves. The empirical task of improving the content does not.
3inseparable aspects
Logic supplies the rule that makes the transition valid. Content supplies the premises and conclusion. The agent or operator carries the transition out. Within the class of inference episodes, none can be removed while the episode is retained.
Rule and content without an operator leave an abstract relation and no episode. Rule and operator without content leave execution with nothing inferred. Content and operator without a rule leave a change of state.
The operator may be a person, an institution, a proof assistant, an algorithm, or a machine. The theorem concerns the function and takes no position on what carries it.
Experience tests their fit to the world, while the inferential form governs the assessment.
Intelligent Epistemology, §5.1
Verdict The order matters more than it looks. MU has a direct proof that uses no later machinery. Applying the method to MU afterwards is therefore not question-begging. Minimality does the rest. Anything strong enough to ground the existence of consistent inference already entails MU. That makes MU the weakest claim sufficient for the task.
▸From the paper
The problem of the criterion untangles once method and subject matter occupy their proper roles.
One of the paper's characteristic moves. The classical problem is not defeated by a stronger argument; it is dissolved by noticing that two of its terms were being asked to do each other's work.
Gödel, Turing, Tarski, Fitch, Arrow, Goodman, No Free Lunch, and algorithmic induction are usually heard as warnings about the reach of reason. The paper reads them as achievements of the same kind: each fixes a domain and classifies exactly what that domain permits. Set beside Shannon's positive classification, they display the principal shapes a full result can take.
In plain words The famous impossibility theorems are not bad news about thinking. They are finished pieces of thinking. Each one takes a precisely stated question and returns the complete answer. Sometimes that answer is that nothing satisfies the question as asked. That is a result. It comes with a proof. Nothing is missing from it.
Suppose a problem asks for an object satisfying a list of conditions. A proof that the admissible class is empty classifies every candidate in that class as inadmissible, so the empty answer space exhausts the requested domain. The proposition is short. Its consequence is not. An impossibility proof is a finished result. A report that returns one has answered the question. Nothing further is owed.
The constructive routes forward are equally explicit. There are four. Revise a condition. Restrict the domain. Enlarge the output type. Add structure. Each changes the problem, and the change is on the record.
What the structure says
Formal ceilings
Gödel's incompleteness theorems require a consequence relation, consistency conditions, and conclusions that follow from premises before they can be stated at all. MU is prior in logical role, because the incompleteness theorem is itself a valid result from its grounds. The domain-specific hypotheses belong to the Gödelian problem, and the complete result is the theory's internal proof closure while the expressible truths may extend past it.
Turing identifies an empty class of universal halting deciders within a model of computation. Tarski returns a level distinction rather than an emptiness: full truth for a sufficiently expressive object language lives in a metalanguage or a restricted hierarchy. Gödel marks a ceiling inside formal reasoning while presupposing the floor on which the ceiling is proved.
What the structure says
Knowability, aggregation, and generalisation
Fitch's paradox isolates a different limit. Suppose every truth is knowable. Suppose also that some truth p is unknown. Then p together with the fact that p is unknown is true, so universal knowability makes it possible to know that conjunction — which yields both that p is known and, by factivity of the second conjunct, that it is not. Under those modal rules, every truth being knowable means every truth is known.
The result classifies the unrestricted schema and names its exact revision points: the schema, the modal logic, or the conception of knowledge. Arrow does the same for aggregation, returning an empty non-dictatorial class for the full package and making the price of each relaxation visible.
Goodman's riddle and No Free Lunch meet at one lesson. Successful generalisation requires structure supplied by the problem. Inductive bias is that structure. Where representation, problem distribution, or invariance is unresolved, the full answer keeps the surviving family.
Algorithmic induction closes the section by making the relativity explicit. A universal machine supplies the coding against which simplicity is measured, and the invariance theorem bounds variation across suitable machines while preserving finite-scale reference dependence. Universal means universal relative to a chosen coding scheme. That qualification returns in Part V, where an adversarial choice of machine turns a theorem into a warning.
Nine classified results and the shape of the answer each returns (§6)
Nine classified results and the shape of the answer each returns (§6)
Result
What the problem gives
Form of the complete answer
Gödel
An expressive, effectively axiomatised theory with arithmetisation and the relevant consistency hypothesis
an internal proof ceiling: the system's own consistency lies beyond its theorem closure
Turing
A model of effective computation and the demand for a total halting decider
the total halting-decider class is empty
Tarski
An expressive object language and the demand for an internally definable truth predicate
full truth moves to a metalanguage or a restricted object language
Fitch
Factive knowledge, modal closure, and universal knowability
universal knowability collapses to universal knowledge
Arrow
Preference orderings, three or more alternatives, and the fairness and independence package
every aggregation rule satisfying the full package is dictatorial
Goodman
Evidence, a projective question, and an unresolved predicate or representation language
a family of projective rules, until representation and invariance are supplied
No Free Lunch
An unrestricted problem class averaged uniformly over all objective functions
uniformly averaged performance is equal; problem structure creates advantage
Algorithmic induction
A coding language or universal reference machine
universality and simplicity are fixed only relative to that reference
Shannon
A source, channel, output alphabet, and the continuity and composition conditions
entropy and channel quantities in their classified information-theoretic form
Read the middle column first. Each row's ceiling exists because its problem was fully stated. The last row shows the same machinery returning a positive classification rather than an empty class.
∅ + certificatea complete classification
A proof that the admissible class is empty classifies every candidate in the defined class as inadmissible. The empty answer space exhausts the requested domain.
Factive knowledge, ordinary modal closure, and the claim that every truth is knowable. If some truth p is unknown, knowing p ∧ ¬Kp would yield both Kp and ¬Kp. Universal knowability collapses into universal knowledge.
A revised condition, a restricted domain, an enlarged output type, or added structure. Arrow's theorem makes the price of each visible rather than hiding it inside a design failure.
The limitation of reasoning is still something known by reasoning.
Intelligent Epistemology, §6.5
Verdict Every ceiling in this table was proved by reasoning, inside a problem that had to be fully specified before the proof could run. That is the paper's structural point about limits: a limitation of reasoning is still something known by reasoning. The certificate that establishes it is a complete answer rather than the absence of one.
▸From the paper
More generally, MU places every limit theorem inside the result supported by its problem: a certified empty class can itself be the complete answer. The limitation of reasoning is still something known by reasoning.
The italics are the paper's. The sentence closes Part II and sets up the whole of Part III, where the same move produces four positive classifications instead of a ceiling.
Part III
The Mathematics of Rational Belief
Part III
Twelve sections of the paper, §7–§18, in six here. The four laws take one section each; full answers folds §13–§15 and channels folds §16–§18.
A wise man proportions his belief to the evidence.
David Hume, An Enquiry Concerning Human Understanding
Film · 2
provedOne rule, four lawsThe generated sequence assembling: scalar plausibility to probability, probability to relative entropy, relative entropy to the two projections, value and cost to choice.≈ 24s
One rule, four laws. The generated sequence assembling: scalar plausibility to probability, probability with hierarchical decomposition to relative entropy, relative entropy with a reference to starting belief and with retention to revision, value and cost to choice (Thm 8.1, Fig. 2).
The theorem states the four arithmetic forms and defers every proof to
the sections that follow: The next four sections prove items
(i)–(iv). Each clause below names the form, the freedom it leaves,
and the theorem that settles it.
Before any arithmetic, the paper draws three lines. Constraint against conclusion. Internal support against connection to truth. Constitutive against hypothetical. The three cut the classical problems at their joints, and §7.5 runs the cut on induction, Goodman, Gettier, and scepticism in a single paragraph each. Theorem 8.1 then announces what the next four sections prove.
In plain words Three distinctions do most of the philosophical work. What you were given versus what follows from it. Whether your reasoning was any good versus whether it happened to land on the truth. Whether a rule is part of what an activity is, or a claim made inside that activity. Get these apart and the famous problems stop overlapping.
Declare what the problem contains. Then make a claim about it. The
three lamps are the same requirement seen three ways. Every one of
them compares the claim to an answer recomputed from the declarations
rather than to an answer key.
content · selection · calculusfidelity to grounds, in three affirmative forms
the problem declares
and the answer claims to
all three fidelities hold · 3 supported classes, 3 reported · nothing smuggled
content fidelityholdsThe report covers all 3 classes the grounds support, so nothing is asserted beyond what the premises carry.
selection fidelityholdsNo point choice was made, so there is no selector to track.
calculus fidelityholdsNo operation was used beyond reading the surviving set.
classes supported3by the grounds as declared
classes reported3by the claim as made
distinctions dropped0supported, and not reported
what the claim returns3 classes: H₁ · H₂ · H₃computed on the survivors, not looked up
Nothing smuggled. Every proposition the report makes is carried by the
declared grounds, no point was chosen that the problem does not fix, and no
operation was used whose structure the problem does not supply.
at machine scale
§28.3 gives this failure its other name. A generative system that reports
content its constraints never paid for instantiates exactly what
Non-Smuggling forbids. The standard applies to machine reporters
unchanged. Hallucination is this violation under thin constraints.
None of the three lamps is a test of effort or of good faith. Each one asks
the same question about a different part of the report: does this piece of
the answer trace back to something the problem contains? Declaring the
missing structure turns every red lamp green without changing a single
number, which is the point. The problem is not that the answer was wrong.
It is that it was not paid for.
proved the three affirmative forms are
Corollary 3.5; the licensing of a compression by a declared selector is
Theorem 3.2 with Corollary 3.3, and the dependence on grounds is Theorem
3.1 (§3). The machine-scale reading is §28.3.
argued the six candidate states,
the stated constraint, the order, and the conjunction A ∩ B are display
choices. The audit itself declares nothing: the honest answer is recomputed
from the grounds on every click. The lamps compare the claim to that
rather than to a stored verdict.
The three fidelities of Corollary 3.5, run as a check. Move the stated grounds on the left and watch each lamp: content fidelity fails when a proposition claims support its premises never carried, selection fidelity when a point is chosen with no selector in the problem, calculus fidelity when an operation uses structure the problem never supplied.
The turn
Constraint against conclusion
A constraint is part of what the problem gives. A conclusion is what those constraints support under the chosen consequence relation. Non-Smuggling governs the passage from the first to the second. Both directions of error are live.
Constraints include more than observations. Far more. Logical relations, symmetries, apparatus, form of answer, and equivalence all qualify. Treating one of them as hidden background makes the problem look more determinate than it is. Treating a conclusion as an input makes a derivation circular. The constraint language of Definition 7.2 lists the five forms the paper uses and requires anything further to be declared before it may affect an answer.
What the structure says
Internal support against connection to truth
The internal question is whether an answer follows from the agent's grounds. The external question is whether those grounds represent the world and arrive through channels that reliably track the relevant truth. They are separate coordinates. Every classical case that trades on the gap between them is exploiting that separation.
A false model can support impeccable updating. A lucky guess can be true without justification. Neither is odd. Neither observation is paradoxical once the two coordinates are drawn apart. Part III's channel and convergence sections classify exactly when internal updating tracks the world, and Part IV's treatment of Gettier is this distinction with a third coordinate added.
What the structure says
Constitutive against hypothetical
A condition is constitutive when it belongs to what the practice is — that an inferential conclusion depend on its grounds, for instance. A claim is hypothetical when it is one proposition assessed within that practice. Keeping the levels distinct is what stops the regress demands from biting.
A model's connection between past and future is a substantive question with an empirical answer. An audit of all inference is itself an inference and therefore stays inside the same constitutive form. That asymmetry is the diagnosis, and §7.5 applies it four times over: induction separates constitutive support from projective hypotheses; Goodman separates constraints from representation; Gettier separates internal support from the external route; scepticism separates global denial from local channel doubt.
The turn
What the four laws claim
Each quantitative domain names an object: plausibility, informational change, starting belief, revision, choice, or learning from a channel. MU generates the laws of each by preserving every dependence the grounds carry, every distinction the state carries, and every equivalence the descriptions carry. Four recurring movements do the work — dependence, retention, refinement, completion — and their arithmetic forms are probability, relative entropy, maximum entropy, and KL projection.
Theorem 8.1 states all four and proves none of them. Its proof is one line long and points forward. The next four sections do the work. The theorem's real content is the uniqueness claim and the exact freedom it leaves. Probability is fixed up to order-preserving regraduation. Relative entropy is fixed up to positive scale. Nothing else is free at all.
The paper is careful about what a uniqueness result means here. Logicians call it relative categoricity; the plainer phrase, which the paper prefers, is uniqueness within the domain. Change the logic, the representation, the reference, the value, the cost, or the output question and a new problem has been posed. It gets its own full answer. That answer may look nothing like the one before it.
4quantitative laws
Probability for determinate Boolean scalar plausibility. Relative entropy for total informational change. Relative-entropy projection for least-assuming completion, with maximum entropy as its finite counting-symmetry form. KL projection for revision. Each is unique in its domain.
Logical and partition relations; expectation or moment constraints; interval and inequality constraints; symmetries and invariances; likelihood or channel conditions supplied by the apparatus. Any further constraint has to be made explicit before it can affect an answer.
A point answer occurs when one answer remains after the relevant equivalences and explicit selectors. A set-valued answer occurs when several inequivalent possibilities remain. The family is then the full answer.
A fully stated domain can admit one mathematical form. Logicians call this relative categoricity; the plainer phrase uniqueness within the domain captures the result. The domain names the object. MU fixes the form that follows from it.
Intelligent Epistemology, Part III opening
Verdict The theorem announcing the four laws proves nothing itself; it defers to the four sections that follow. What it does supply is the shape of the claim. Each law is unique in its domain, up to a stated freedom — an order-preserving regraduation, or a positive scale — and Full-Answer Preservation carries every remaining plurality as a family and every empty or unattained case as a certificate.
▸From the paper
Four recurring movements do the work: dependence, retention, refinement, and completion. Their arithmetic forms are probability, relative entropy, maximum entropy, and KL projection.
The sentence that turns Part I's philosophy into Part III's mathematics. Each movement is a way of not adding and not discarding. Each arithmetic form is what that discipline becomes once a domain has said what it is about.
Name the object: a Boolean algebra of propositions, a total plausibility order on a rich scalar range, a scalar state complete for conjunction and disjoint union, and symmetric refinement. Three conditions. From them MU generates five clauses of a calculus, and from those five the product rule, the sum rule, and complementation follow with one freedom left over.
In plain words Nobody has to decide that beliefs obey the rules of probability. If you are grading propositions on a single number, that number behaves properly under and, or, and not, and you can always split a case into equally likely parts, then the ordinary rules of probability are the only thing left. Everything else has been ruled out.
The forcing chain. Three domain conditions on the left, the five clauses MU generates from them in the middle, the three rules of the calculus on the right. Each arrow is a step in the proof, and the dashed return marks the one freedom left: order-preserving regraduation.
The turn
Naming the object
Boolean and scalar name an object here, the way group names an algebraic object. A determinate Boolean plausibility domain has three conditions. First, a Boolean algebra of propositions with a total plausibility order whose scalar range is separable and operationally rich. Second, a scalar state complete for conjunction and disjoint union, meaning the scalar inputs and the Boolean relations exhaust the operation-relevant information. Third, symmetric refinement: the domain can be cut into equal parts, and the cutting behaves. Rational subdivisions exist. A common refinement of two descriptions preserves the represented event. The realised partial ranges are dense in their target intervals.
Notice what is not assumed. No betting interpretation. No scoring rule. No axioms about preference over acts. None of it. The conditions describe what the domain is. MU supplies every inferential law of its calculus.
What the structure says
Five clauses before any rule
Completeness of the scalar state plus dependence on grounds give scalar substitution: equal scalar states are interchangeable. That yields universal operations for conjunction and disjoint union. Full-answer preservation gives retention of support. Each operation is then strictly increasing on every nondegenerate argument whose order distinction the state preserves. Symmetric refinement and presentation invariance give refinement coherence.
Boolean coherence follows from logic alone: expressions that are logically identical represent one proposition and therefore carry one value, which delivers associativity, identity, zero, complementation, and distributivity in a single stroke. Density then fills every interval, so the operations are continuous, and a rectangle squeeze upgrades separate continuity to joint continuity.
None of the five mentions probability. Not once. They are what the object has to satisfy before any rule can be written. The rules are what remains once all five are in force.
What the structure says
The two representations
Conjunction first, then disjoint union. A continuous, associative, strictly increasing operation with the top as identity and the bottom as annihilator can be regraduated into multiplication. Symmetric refinement then additivises disjoint union on the same scale.
One freedom survives that pair: the scale could be any positive power. Boolean distributivity removes it. A proposition and its negation are disjoint and exhaust the algebra, so they sum to one, and inclusion–exclusion gives the general sum rule. Both constructions are set out in the technical block below.
The turn
The neighbours
Finite qualitative systems show the boundary. Kraft, Pratt, and Seidenberg exhibit a five-atom total order that no probability measure agrees with, and Scott supplies the necessary and sufficient cancellation conditions. That domain returns an empty representation class. By Proposition 6.1 the empty class is its complete answer. Scalar completeness, symmetric refinement, substitution, and common-refinement coherence are what select the probability domain instead.
Partial orders, vector-valued states, and context-sensitive operations get their own full answers under their own specifications. So does the non-Boolean case: projection lattices in Hilbert space are classified by Gleason-type theorems under their own dimensional, additive, and regularity conditions. The paper does not claim probability everywhere. It claims probability exactly where the domain is this one. The difference is the point.
Three remarks close the section by locating rival derivations rather than dismissing them. Dutch-book arguments supply an operational consequence of incoherence under a betting interpretation. Accuracy-first epistemology derives probabilism from dominance with respect to strictly proper scoring rules. Preference-first routes recover probability and utility together from axioms on choice. Each begins from different grounds and converges on overlapping quantitative forms. The paper marks one question as still open: whether accuracy dominance and dependence on grounds force one another.
▸Technical · §9.2 · §9.3
What this block carriesThe associative-representation argument that turns conjunction into multiplication, and the additive one that turns disjoint union into addition.
Monotone Cauchy
h(x + y) = h(x) + h(y) ⟹ h(x) = h(1)·x
A monotone additive function on the positive reals is linear. Additivity settles the positive rationals; rational sequences rising and falling to a real then squeeze the value by monotonicity.
If doubling the input doubles the output and the function never doubles back, the function is a straight line through the origin.
Fix an interior point a and form its integer powers under the operation. They decrease to the bottom of the interval, each has a unique nth root, and the rational powers turn out order-dense. Define λ as the supremum of rationals whose power still dominates x. Order density makes λ a continuous strictly decreasing bijection, additivity of λ over the operation follows on rational powers and extends by continuity, and g = e−λ carries the operation to ordinary multiplication.
Keep applying the operation to one fixed value and you get a ruler. The ruler turns out to be a logarithmic one, so on its scale the operation is multiplication.
Symmetric refinement supplies, for every n, a unique equal part that n copies of exhaust the whole. The resulting rational points are order-dense and additive. Rescaling by their supremum makes disjoint union addition. Boolean distributivity then forces the rescaling to be a power, and the power is absorbed.
Cut the certain event into n equal pieces, count pieces, and you have a scale on which disjoint alternatives add. Distributivity then removes the last freedom in the scale.
A monotone one-variable section can only fail continuity by a jump, and a jump omits an open interval from the realised range. Density excludes that. A four-corner rectangle squeeze upgrades separate continuity to joint continuity.
If every value in between is reached somewhere, the operation cannot skip.
Transcribed from the paper, which marks every one of these proofs
optional on a first reading — and prints them anyway.
The full text ↗
3domain conditions
A Boolean algebra with a total plausibility order on a separable, operationally rich scalar range. A scalar state complete for conjunction and disjoint union. Symmetric refinement, with the realised partial ranges dense in their target intervals.
Scalar substitution, retention of support, refinement coherence, Boolean coherence, continuity. MU generates all five from the three domain conditions before any rule of probability is written down.
The product rule for conjunction, inclusion–exclusion for disjunction, and complementation to one. Probability is MU's unique scalar Boolean calculus, up to order-preserving regraduation.
Kraft, Pratt, and Seidenberg exhibit a five-atom total qualitative order that admits no agreeing probability measure. That neighbouring domain returns an empty representation class. Scott supplies the cancellation conditions separating the cases.
Accuracy-first epistemology derives probabilism from dominance under strictly proper scoring rules. Preference-first routes recover probability and utility together from axioms on choice. The route here begins from dependence on grounds and the domain's richness conditions. The three converge on overlapping quantitative forms, and whether accuracy dominance and dependence on grounds force one another remains open.
Probability is Epistemic Zero written as arithmetic: every supported distinction is preserved and every undetermined distinction remains open.
Intelligent Epistemology, §9
Verdict The uniqueness claim is exact and worth stating carefully. It is not that one function is picked out; it is that every representation admitted by the domain is a monotone relabelling of a single calculus. Weaken the domain and the neighbouring cases are real: a five-atom qualitative order with no agreeing measure returns an empty representation class. A partial comparison returns the family of compatible representations.
▸From the paper
Probability is Epistemic Zero written as arithmetic: every supported distinction is preserved and every undetermined distinction remains open.
The sentence explains why the derivation needs both halves of the bilateral rule. Preservation of distinctions gives strict monotonicity; adding nothing beyond the constraints is what leaves the calculus with exactly one degree of freedom.
▸The proof idea · Thm 9.2argued
MU forces probability
HypothesesA determinate Boolean plausibility domain: Boolean algebra with a total order on an operationally rich scalar range, a scalar state complete for conjunction and disjoint union, and symmetric refinement with dense realised ranges.
Completeness of the scalar state plus dependence on grounds make every represented operation factor through the scalar values.
Full-answer preservation makes each operation strictly increasing wherever the state preserves an order distinction.
Dense refinement fills every interval, so the operations are continuous.
The associative-representation argument then regraduates conjunction into multiplication.
Symmetric refinement builds rational parts of the certain event, which additivise disjoint union; distributivity forces the residual freedom to a power, which is absorbed.
Complementation follows because a proposition and its negation are disjoint and exhaust the algebra.
One transition between probability states admits two descriptions: a flat joint account, and a staged account that reports a marginal change and then conditional changes on each branch. MU requires both to carry the same total. That single demand, plus the requirement that conditional branches be weighted by the revised state, forces relative entropy up to a positive scale.
In plain words You can describe a change all at once or in stages. If both descriptions are of the same change, they have to add up to the same amount. Insist on that and there is only one way to measure informational change. It is the one information theory already uses.
One transition, two descriptions. Sum the joint cell by cell, or stage
it as a marginal move plus what each branch does afterwards. MU asks
the two totals to agree, and everything else in this section follows
from that demand.
D(Q‖P) = D(QX‖PX) + Σx QX(x) D(QY|x‖PY|x)branch weights come from the revised state, never the old one
flat 0.736463 · staged 0.736463 · disagreement zero to machine precision · coarse-grained 0.451816
run it to
The account, both ways. The right-hand column is summed over all 6 joint
cells with no staging; the left is built one term at a time.
the staged accountweightconditionalcontributes
marginal move on X · D(QX‖PX)——0.049857
branch x₁ · D(QY|x₁‖PY|x₁)0.75000.9154750.686607
branch x₂ · D(QY|x₂‖PY|x₂)0.25000.0000000.000000
staged total——0.736463
flat total · summed over the 6 joint cells——0.736463
In nats the two accounts disagree by zero to machine precision. Both totals are
computed from the state on screen: the flat one sums 6 cells, the
staged one adds three terms, and neither is derived from the other.
Branch weights are QX, the revised state, as Lemma 10.1
requires. Switch them for the old state and the two accounts stop agreeing.
Coarse-graining, Lemma 10.4: push both states through one kernel that merges y₂ and y₃
D(Q‖P)0.7365before the kernel
D(QK‖PK)0.4518after it
the fall0.2846never negative, and here is why
the remainder0.2846the proof's own nonnegative term
The lemma decomposes one joint two ways. Through the merged cell it
collapses; through the fine cells it carries a surplus that is a sum of
divergences and cannot be negative. The fall and that surplus are
computed separately here. In nats they differ by zero to machine precision.
The staged account reads 0.736463 nats and the flat one reads 0.736463. Coarse-graining then takes the total down to 0.451816, a fall of 0.284648, and that fall is exactly the information the merged cell no longer distinguishes.
The chain rule is not a convenience. It is what picks relative entropy out
of the wider family of monotone divergences: drop the demand that the
staged and flat descriptions balance exactly. Rényi and the rest come back in. Everything downstream of this card — the two projections, the
Pythagorean split, the convergence argument — rests on the disagreement
row above.
proved the accounting identity is
Theorem 10.2 (I4), the final-state weight is Lemma 10.1, and the
coarse-graining inequality is Lemma 10.4 (§10).
argued the reference joint, the
three signal cells, and the tilt direction g = (−1, 0, +1) are display
choices. The identities hold on every finite informational-change
domain. Every figure on this card is computed for the state shown rather
than fitted to it.
▸The measured residuals
flat against staged
2.220e-16
the fall against the surplus
1.665e-16
Both totals are summed from the state on screen and neither is derived
from the other. A figure of 0 means the two landed on the same double.
Anything else is the width of one. Where a line above reports how far two
accounts disagree, this is the figure it stands for.
The chain rule as a ledger. Drag a conditional branch and watch the staged account fill: a marginal term, then one conditional term per branch weighted by the revised state. The flat total on the right has to match. Push both states through the coarse-graining and the total can only fall (Thm 10.2, Lemma 10.4).
The turn
The object, and the direction
An informational-change domain represents a transition by a nonnegative scalar magnitude with the unchanged transition as its identity. It admits independent product composition, disjoint branch composition, symmetric refinement, and equivalent flat and hierarchical descriptions. Its realised magnitudes are rich enough for the continuous representation of those compositions.
Before anything else the direction has to be fixed. Lemma 10.1 fixes it. A conditional change of magnitude d on a final-state branch of probability q contributes exactly qd to the total. The proof refines the final-state partition into equal copies, uses symmetry and branch additivity on rationals, and extends by continuity. The consequence is that conditional changes are averaged with the revised marginal rather than the old one, and reversing the transition exchanges both the arguments and the branch weights.
What the structure says
Six properties, and the one that matters
The identity law and full-answer preservation give nonnegativity with strict zero only at agreement. Dependence on grounds gives invariance under common relabelling, since a relabelling preserves the transition it describes. Independent components compose by a universal scalar operation, and the associative-representation argument that ran the probability derivation now supplies a scale on which that composition is addition.
Then the chain rule arrives. A joint transition has a flat description and a staged one, and both represent the same transition, so dependence on grounds forces them equal. With final-state weighting already fixed, the equality is the exact marginal–conditional decomposition. Information monotonicity is a corollary rather than an axiom: decompose one joint two ways and the inequality falls out.
What the structure says
Pinning the constant
The characterisation proof is a chain of three moves. Write c(n) for the divergence of a point mass from the uniform on n outcomes. Splitting n into blocks makes the chain rule give additivity across products. A carefully built stochastic kernel maps the uniform on n+1 outcomes onto the uniform on n while fixing the point mass, so monotonicity gives c(n+1) ≥ c(n). Additive and monotone means logarithmic. That is the first move.
Next, refinement leaves the value alone, padding with commonly empty cells changes nothing, and a uniform nested inside a larger uniform contributes the logarithm of the ratio. Finally, build a joint whose rows are uniform blocks sized so that the first argument nests inside the second, evaluate it through the row variable and through the permuted whole, and equate. The KL formula appears on strictly positive rational pairs, and continuity extends it everywhere.
One positive constant survives. It is a unit rather than a law. Corollary 10.6 then extends the formula across the support boundary by lower semicontinuity, which is where the infinity comes from.
The turn
One divergence, two uses
The section closes with a remark that organises the next two. Prior completion and updating use one classified divergence with different reference objects. The first reference is a measure permitted by the representation and apparatus. The second is the current state of belief. The same divergence governs both maps.
That is not an economy of notation. It is the reason the two projections have the same theory: least-assuming completion and retention are the same instruction pointed at different objects. Carry forward every feature compatible with the relevant constraints. Change the rest.
▸Technical · §10
What this block carriesThe characterisation proof: uniform divergences grow logarithmically, refinement leaves the value alone, and two evaluations of one nested pair pin the constant.
Uniform divergences and logarithmic growth
c(mk) = c(m) + c(k)
c(n) = κ·log n, κ = c(2) > 0
Write c(n) for the divergence of a point mass from the uniform on n outcomes. Splitting n = mk into m blocks of k makes the chain rule give c(mk) = c(m) + c(k). A carefully built stochastic kernel maps the uniform on n+1 outcomes onto the uniform on n while fixing the point mass, so monotonicity gives c(n+1) ≥ c(n). Additive and monotone forces logarithmic.
Spreading each cell uniformly over a block leaves the divergence unchanged, because the conditional terms are uniform against uniform and vanish. Adjoining cells both states leave empty changes nothing. A uniform on s cells nested inside a uniform on t contributes κ·log(t/s).
Build a joint whose rows are uniform blocks of sizes chosen so that the first argument nests inside the second. Evaluate through the row variable and through the permuted whole. Equating the two gives the identity on strictly positive rational pairs, and continuity extends it.
The lower-semicontinuous envelope on a finite closed simplex. Where the revised state is dominated, approach every zero coordinate through strictly positive pairs. Where it is not, every approximating sequence carries a term diverging upward.
Transcribed from the paper, which marks every one of these proofs
optional on a first reading — and prints them anyway.
The full text ↗
6generated properties
Identity with strict retention; invariance under common relabelling; additive composition of independent components; exact marginal–conditional decomposition; contraction under every common stochastic channel; continuity on positive simplex interiors.
Push both states through the same stochastic kernel and the measured change can only fall. The inequality is derived from exact accounting rather than assumed alongside it.
The lower-semicontinuous envelope of the positive-simplex formula. Where the revised state puts mass on a cell the old state leaves empty, the accounting does not close.
Relative entropy is Epistemic Zero's ledger: every change is counted once, and every equivalent description balances to the same total.
Intelligent Epistemology, §10
Verdict Exact conditional decomposition is what selects KL from the wider family of monotone divergences. Weaken the chain rule to monotonicity alone and the Rényi divergences walk back in. The paper's own price list in §18 makes the trade explicit: relative entropy costs you the exact chain rule and nothing else.
▸From the paper
Relative entropy is Epistemic Zero's ledger: every change is counted once, and every equivalent description balances to the same total.
Counted once is the operative phrase. The uniqueness of KL is not a claim about convenience or tradition; it is what survives when double-counting and under-counting are both ruled out.
▸The proof idea · Thm 10.5argued
MU forces relative entropy
HypothesesA finite informational-change domain: a nonnegative magnitude for each transition, independent product composition, disjoint branch composition, symmetric refinement, and equivalent flat and hierarchical descriptions.
Final-state branch weighting fixes the direction: a conditional change on a branch of revised probability q contributes q times that change.
Because one transition has both a flat and a staged description, dependence on grounds forces the exact marginal–conditional chain rule.
The chain rule applied to a point mass against the uniform makes the uniform divergences additive across block splittings.
A stochastic kernel carrying the uniform on n+1 outcomes to the uniform on n, while fixing the point mass, makes them monotone.
Additive plus monotone gives logarithmic growth with a positive constant.
Nesting one uniform inside another and evaluating a single joint two ways gives the KL formula on strictly positive rational pairs; continuity finishes it.
Starting belief is MU applied before evidence arrives; updating is MU applied through time. Both are the same instruction — carry forward every feature the constraints permit — pointed at different objects. Projecting from a permitted reference gives least-assuming completion, with maximum entropy as its finite counting-symmetry form. Projecting from the current state gives KL revision.
In plain words Two questions look different and turn out to be one. Where should you start before you know anything? Where should you move when you learn something? Both answers are: go to the nearest state that satisfies the constraints, where nearest is measured by the ledger from the last section. Only the starting point changes.
The same minimisation runs twice. Completion measures from the
reference the problem permits; revision measures from the state
belief is already in. They land in different places, and neither
landing is a choice anyone made.
μ →𝒞 P₀ · P →𝒞 P₁argminQ ∈ 𝒞 DKL(Q‖·), both times
completion D(P₀‖μ) = 0.0302 · revision D(P₁‖P) = 0.3864 · the two landings sit 0.0264 apart in total variation
D(P₀‖μ)0.0302completion, from the permitted reference
D(P₁‖P)0.3864revision, from the current state
H(P₀)1.0684the completion's entropy, in nats
apart0.0264total variation between the two landings
Theorem 11.3, both sides on the completion point: DKL(P₀‖U) reads
0.030229 summed over the three atoms, and log 3 − H(P₀) reads
0.030229 from the entropy beside it. The two disagree by
zero to machine precision. Least-assuming completion and maximum entropy are
one problem written in two coordinate systems. This is the row that says so.
Theorem 12.6, both sides: D(Q‖P) = D(Q‖P₁) + D(P₁‖P), for every Q in 𝒞₁
the splitD(Q‖P)D(Q‖P₁)D(P₁‖P)
summed over the three atoms, each term on its own0.4112930.0249150.386378
The left side reads 0.411293. The two right-hand terms add to
0.411293. They disagree by zero to machine precision. Move Q anywhere
along the family and the split holds. The angle it names is informational,
not the angle the drawing appears to make.
Corollary 12.8: alternate between 𝒞₁ and a second family, and the walk converges on their intersection
steps to the tolerance: —stopped by: not run
Not run yet. The walk starts at the current state, projects onto 𝒞₁,
projects that onto 𝒞₂, and repeats. Each step is a KL projection. The corollary
says the sequence reaches the projection onto the intersection.
Remark 10.2 is one sentence and it decides a great deal. Completion asks
what the least-assuming state compatible with the constraints is, measured
from what the problem permits before belief arrives. Revision asks what the
least-assuming state compatible with the constraints is, measured from
where belief already stands. Feed the same constraint to both and they part
company, which is why a prior and an update are different objects governed
by one law.
proved completion is Theorem 11.1 with
Theorem 11.3, revision is Theorem 12.1 with Theorem 12.2, the split is
Theorem 12.6, and the alternation is Corollary 12.8 with Lemma 12.7
(§11–§12).
argued three atoms, the
observable f = (0, 1, 2), and the second family p₂ = d are display choices;
the results hold on any finite outcome space with a linear family carrying
a strictly positive point. Projections are solved by bisection on the
exponential-family parameter to a stated tolerance, and the alternating run
carries a named clamp at 200 steps, reported whenever it is what
stopped the walk.
▸The measured residuals
divergence against entropy
3.331e-16
the Pythagorean split
2.220e-16
Each of these is one quantity computed two ways from the state on
screen, with neither route derived from the other. A figure of 0 means
the two routes landed on the same double. Anything else is the width of
one. Where a line above reports how far two readings disagree, this is
the figure it stands for.
The stated tolerance the bisection stops at is 1e-12,
and the alternating run reports which of the two ended it.
One divergence, two reference objects. Drop a constraint set onto the simplex and watch both projections run. Completion starts from the permitted reference μ. Revision starts from the current state P. The right angle is the Pythagorean identity; the alternating path between two constraint families is Corollary 12.8 converging.
The turn
What a permitted reference is
Before evidence, the problem still supplies something: the comparison structure against which any concentration will be measured. Definition 11.1 makes it precise. A positive measure class is permitted when it belongs to an invariant assignment across the representations the problem and apparatus allow, with the transformations that preserve the statistical question carrying one presentation's class to another's. The permitted family is every such value, taken up to positive scale. No member is preferred.
MU then selects, for each permitted reference, the feasible states carrying the least informational departure from it. One reference and one attained minimiser give a point prior. Several references or several minimisers give the complete attained family. Every unattained permitted case appears separately with its attainment boundary.
What the structure says
Maximum entropy in counting coordinates
On a finite atom space with a permutation-symmetric counting reference, the relative entropy to the uniform is log N minus the Shannon entropy. The two optimisation problems therefore have identical solutions. Maximum entropy is not a separate principle. It is least-assuming completion written in counting coordinates.
With the reference fixed and the feasible set determined by attained expectation constraints, an interior solution takes exponential-family form. That form is the Euler–Lagrange equation for a strictly convex objective under affine constraints. Declare a complexity function together with a moment constraint on it. The same theorem returns a Boltzmann prior over hypotheses. That is an Occam ordering relative to the chosen code. Change the code and the case changes with it. Say which code.
What the structure says
Where the reference comes from, and where it runs out
Problem-preserving transformations determine the reference answer, and three singleton cases are named. A compact transitive symmetry with a unique normalised invariant measure. A problem that supplies the Fisher information metric and asks for its induced volume class. An affine coordinate with congruence invariance, which gives Lebesgue and, on a bounded interval, its normalised uniform case.
Noncompact symmetry is where the machinery stops short. The paper says so precisely. Translation invariance gives every unit interval the same mass, and countable additivity over a disjoint cover of the line then forces total mass zero or infinite. Scale invariance runs the same argument on dyadic intervals. The invariant class survives. The normalisation does not. A proper probability emerges only when the problem adds location, scale, truncation, likelihood, or another normalising structure.
The turn
Revision as retention
Updating applies the same rule through time. The current state carries the information already earned; new constraints purchase a specific departure from it. Revision carries forward every feature compatible with the constraints and changes exactly what joint satisfaction requires. The complete update is therefore the whole set of divergence minimisers over the feasible set.
Ordinary conditioning is the special case where the constraint sets an event's probability to one. Bayes' rule appears there as the minimiser. Jeffrey conditioning is the case where it sets that probability to some other value: the old conditionals inside and outside the event survive, and only their mixture weights move. Both are consequences here rather than additional postulates. Neither was assumed.
On a nonempty closed convex feasible set with a point of finite divergence, strict convexity gives existence and uniqueness. Alternating projections between two linear families converge to the projection onto the intersection, which the Pythagorean identity and Pinsker's inequality prove together. The telescoped identity makes the step divergences summable. Pinsker turns that into vanishing total-variation steps. Compactness supplies the limit.
Path independence closes the section. For compatible finite linear constraints, the simultaneous projection is a function of the joint feasible set. That set is invariant under the order the constraints are presented in. Alternating projection converges to it. A single sequential pass reaches the same point when the projection operators commute or preserve one another's constraint families, and the paper is careful not to claim more: nonlinear and nonconvex constraints need their own algorithms and their own convergence theorems.
▸Technical · §11 · §12.2
What this block carriesExponential-family form, the Pythagorean identity, Pinsker, and the alternating-projection convergence proof.
Exponential-family form
dP*/dμ (x) = exp(Σⱼ λⱼ·fⱼ(x)) / Z(λ)
The Euler–Lagrange equation for the strictly convex relative-entropy objective under affine expectation constraints. Existence and boundary cases still need the usual integrability and attainment hypotheses.
On a linear family with a strictly positive interior point, the Lagrange equations make the log-density ratio at the projection an affine function of the constraint statistics. Its expectation is therefore the same for every feasible state, and expanding the divergence through the projection splits it exactly.
Coarse-grain by the indicator of where the first state exceeds the second and apply information monotonicity. The binary case reduces to a function with value and first derivative zero at agreement and a nonnegative second derivative.
The Pythagorean identity telescopes, so the step divergences are summable and Pinsker sends the total-variation steps to zero. Compactness supplies an accumulation point; even and odd subsequences lie in the two closed families and their asymptotic equality puts the limit in the intersection. Strict convexity makes it unique, so the whole sequence converges.
Noncompact symmetry and the normalisation obstruction
dx under translation · dx/x under rescaling
Translation invariance gives every unit interval the same mass. Countable additivity over a disjoint cover of the line then forces total mass zero or infinite. Scale invariance runs the same argument on dyadic intervals. The invariant class survives; the normalisation does not.
Transcribed from the paper, which marks every one of these proofs
optional on a first reading — and prints them anyway.
The full text ↗
Four worked examples, one per shape (§11.4)
Four worked examples, one per shape (§11.4)
Case
The structure stated
The complete answer
A point
Six outcomes with the full permutation symmetry of a fair die
uniform, entropy log 6; the symmetry route and the projection route meet at the same point
A family
Two invariant references left standing, with neither preferred by the problem
the indexed family of reference-and-minimiser pairs; selecting one would add a preference the problem never supplied
An improper class
Translation invariance on the real line
the sigma-finite Lebesgue class together with the normalisation obstruction; further structure restores propriety
An unattained optimum
Two atoms, uniform reference, and the open constraint p₁ > ½
the case, the infimum 0, and the boundary point (½, ½) that the constraint excludes
The paper closes each of the four by exhibiting it rather than describing it. That is the section's method in miniature: the shapes of Definition 3.4 are not a taxonomy imposed from outside, they are what the machinery produces when it is run.
argmin DKL(P‖U) = argmax H(P)on a finite atom space
With a permutation-symmetric counting reference, DKL(P‖U) = log N − H(P). Minimising departure from counting and maximising Shannon entropy are the same problem in different coordinates.
Revision carries forward every feature of the current state compatible with the new constraints. It changes exactly what joint satisfaction requires. Every minimiser is retained; every unattained optimum is reported.
For a linear family on a finite outcome space with a strictly positive interior point, the projection splits the divergence exactly. Alternating projections between two such families converge to the projection onto their intersection.
On two atoms with uniform reference, minimise over the open case p₁ > ½. The infimum is 0, approached as p₁ falls to ½. The complete answer is the case, the infimum, and the excluded boundary point (½, ½).
Maximum entropy is Epistemic Zero before evidence: the selected state carries the least concentration compatible with the constraints and the reference.
Intelligent Epistemology, §11
Verdict The retention half is doing real work. A classified divergence fixes a geometry; without the instruction to change no more than the constraints require, any update policy could be paired with it. §18 lists that exact pairing as one of the rivals the retention rule excludes. Where several references or several minimisers survive, the answer is the family — and where the optimum is not attained, the answer is the case, the infimum, and the boundary.
▸From the paper
KL projection is Epistemic Zero through time: each new constraint changes exactly what it touches and carries every compatible feature forward.
Set this beside the maximum-entropy line from §11 and the pair is the whole of Part III's middle. Same instruction, same divergence, different starting object.
▸The proof idea · Thm 12.1argued
MU's retention law and KL projection
HypothesesA current state, a feasible constraint set, and a case that represents informational change by the classified divergence.
New constraints purchase a departure from the current state, and the divergence orders feasible revisions by the size of that departure.
A revision with strictly greater divergence introduces change the constraints did not require, while a feasible revision with less is available.
MU therefore selects the minimisers, and full-answer preservation returns the whole minimiser set rather than one member of it.
On a nonempty closed convex feasible set containing a point of finite divergence, strict convexity gives existence and uniqueness.
Standard conditioning is the special case where the constraint sets the probability of an event to one.
Jeffrey conditioning is the case where it sets that probability to q: the old conditionals inside and outside the event survive, and only their mixture weights move.
definedEvery answer shapeMass on a simplex resolving four ways: a point fixed by symmetry, an orbit its symmetry cannot separate, a family nothing selects among, and a boundary the constraints never reach.≈ 22s
Every answer shape. Mass on a simplex resolving four ways: a point fixed by every symmetry, an orbit its symmetry cannot separate, a family nothing in the problem selects among, and a boundary the constraints approach without reaching. The four settings are the paper's own worked examples (Def. 3.4, §11.4).
Sometimes the answer is a set, and that is the answer
When several inequivalent completions survive, the credal family is the result. A question can demand one number while its grounds determine an interval. The paper treats that as a finding rather than a failure, then shows what has to be added before a point appears: a value, a cost, a selector, or a protocol. Each addition is a new input on the record.
In plain words Some questions do not have one right answer given what you know. Pretending otherwise is not rigour. The honest reply is the whole set of answers your evidence permits, together with the boundary of what it fixes. If you need one number to act on, you have to supply something more. Say what.
Three inputs and no others: what the problem permits, what it rewards,
and what departing costs. Push value as hard as you like at the state
the reference left empty. Then widen the family and watch how many
actions the answer honestly contains.
run τ to
At τ = 1.00 the tilted state sits 0.1693 nats from the
reference, and x₅ holds 0.000000 of the mass.
τ log Z1.5157the value at the maximum, Thm 14.1
Eρ*[V]1.6851value the tilted state earns
τ D(ρ*‖μ)0.1693what it pays to depart
ρ*(x₅)0.000000the state the reference left empty
Theorem 14.1, both sides: Eρ[V] − τD(ρ‖μ) = τ log Z − τD(ρ‖ρ*)
candidate stateleft sideτD(ρ‖ρ*)what it gives up
ρ = ρ*, the tilted state1.5157290.000000sits on the tilt, so it scores τ log Z and no state scores more
ρ = μ, stay at the reference1.3250000.190729earns Eμ[V], pays nothing to depart
ρ = all mass on the best supported state——earns the largest V it can reach, pays log(1/μ) to get there
The right-hand column is summed from the two states rather than read off
the first, so the largest disagreement between the two sides is a
measurement.
Theorem 14.4: the undominated set over the credal family, action by action
a₁ties at its best across two states1.3250dominated
a₂strong where the reference is heaviest1.6100survives
a₃the same payoff whatever happens1.4000dominated
a₄its best payoff sits on x₅0.3350dominated
At width 0 the family holds one state. Every comparison is decided.
The answer is a single action. Widen it and the comparisons stop agreeing
with each other.
Nonemptiness, measured rather than repeated: the width was swept across
200 settings from 0 to 0.180 and the undominated set
counted at each. Smallest count seen: 1. Times it was empty: 0.
State x₅ holds 0.000000 of the tilted mass. Action a₄ puts
its largest payoff there. Neither fact helps it. Value reweights the
possibilities the reference already carries. A possibility
carrying zero mass has nothing to reweight.
A decision answer with several members is not an unfinished decision. It is
the complete one, and it stays complete until an ambiguity rule, a
reference, a loss refinement, or an institution supplies the structure that
narrows it. Reporting a single action from a family that supports several is
a selection with no selector in the problem.
proved the tilt and its identity are
Theorem 14.1, the support statement is Corollary 14.2, the three roles are
Proposition 14.3, the cost path is Proposition 14.5, and the nonempty
undominated set is Theorem 14.4 (§14).
argued the five states, the four
utility vectors, and the shape of the credal family are display choices, not
results. Expected utility is affine in the state. Dominance over the family is
decided at its extreme points. The card checks those. State x₅
is not special-cased anywhere in the arithmetic; it holds zero because the
reference gives it zero and a finite exponential leaves zero alone.
▸The measured residual
the two sides of the tilted score
—
Each of these is one quantity computed two ways from the state on
screen, with neither route derived from the other. A figure of 0 means
the two routes landed on the same double. Anything else is the width of
one. Where a line above reports how far two readings disagree, this is
the figure it stands for.
A credal family, a finite action set, and the two ways to get from one to the other. Move value and the cost of departure and watch the tilt run; watch the undominated set shrink and refuse to empty. The blacked-out region is Corollary 14.2: zero reference mass stays zero however hard the value pushes.
The turn
Plurality has several causes
A credal state is the full set of point-valued answers a problem permits, and it coincides with the raw feasible simplex only when the problem independently supports every feasible point. Plurality can arise from an incomplete plausibility comparison, from several permitted reference structures, or from several tied minimisers. Attainment data stay a separate part of the answer.
Sparse evidence, unresolved representation choices, several natural references, and uncertain channel models each produce a plural state. None of them is a defect in the reasoning. Each is a fact about what the grounds fix. Report it as one.
What the structure says
Uniqueness and permissivism, reconciled
The long dispute between rational uniqueness and epistemic permissivism turns out to concern two different levels. Identical problems determine the same complete permission set. It may contain one member or many. Uniqueness governs the answer as a whole. Permissivism describes its members. Both were right.
Agents sharing a credal family may still act differently, because utilities, losses, and resources belong to the downstream decision problem. Different evidence, representation, reference families, channel models, or computational access define different belief problems in the first place. Full identity of every answer-relevant input yields identity of the full permission set. That is the only uniqueness claim the paper makes.
What the structure says
When one number is required anyway
A reporting or decision protocol may demand a point while the belief state remains a set. Resolute completion supplies one: given a reference with finite divergence to some member, take the divergence-minimising member of the family and hold it fixed until new information changes the problem. On a closed convex family with a finite-divergence point, strict convexity makes it unique. Each permitted reference yields its own resolute case. A one-number protocol becomes determinate only when it also supplies a reference or a selector.
The alternative route chooses an action directly from the credal set. Credal dominance is transitive and irreflexive, so the undominated set is never empty. Every action outside it is worse under every state the problem permits. A singleton fixes the action outright. Several survivors are the complete decision answer until an ambiguity rule, reference, loss refinement, or institutional protocol arrives.
The turn
Choice, and the boundary it inherits
Once a value functional and a cost of departure are supplied, the geometry from §10 fixes the choice distribution outright. Expected value minus the departure cost is maximised at exactly one state. The proof is an exact rewriting rather than a bound. The objective equals a constant minus the divergence to the optimum, so nonnegativity of divergence reads off both the maximiser and its uniqueness.
The corollary that follows is the sharpest line in Part III. The chosen state is equivalent to its reference in both directions. Anything the reference gives zero mass keeps zero mass. A hard exclusion can be represented by restricting the reference, which removes support. Nothing can create it. Ever. Values reweight the possibilities the reference already carries, and a truth, action, or population outside the represented measure class enters only through an enlarged representation.
The two ends of the cost path are worth holding on to. As the cost of departure grows, the chosen state returns to the reference. As it falls to zero, mass concentrates on the maximisers of value inside the support, weighted in proportion to their reference masses. Part V uses exactly this when it prices alignment.
Section 15 gathers the sequence in one plate. Scalar plausibility yields probability. Probability with hierarchical decomposition yields relative entropy. Relative entropy with a permitted reference yields starting belief; with retention it yields revision; with value and cost it yields choice. Each arrow names the structure a domain supplies. Each box is the form MU forces once that structure is present.
▸Technical · §14
What this block carriesThe exact rewriting behind the choice theorem, and the geometry it leaves on a smooth state space.
The identity behind the maximiser
Eρ[V] − τ·DKL(ρ‖μ) = τ·log Z − τ·DKL(ρ‖ρ*)
Because the optimal density is strictly positive almost everywhere, the log-density ratio to it splits into three terms. Integrating against any admissible state gives an exact identity, not a bound. Nonnegativity of relative entropy then reads off the maximiser and its uniqueness in one line.
On a state space carrying a metric and volume form, with a strictly positive smooth reference density and a smooth value, the local score of the tilted state is the sum of a reference term and a value term relative to that geometry.
Under the usual smoothness, nonexplosion, confinement, and ergodicity hypotheses, the overdamped Langevin diffusion is reversible with the tilted law as its stationary measure. Within the isotropic reversible class this is the standard relaxation; preconditioned and nonreversible processes reach the same law by other routes.
Transcribed from the paper, which marks every one of these proofs
optional on a first reading — and prints them anyway.
The full text ↗
3separate inputs
The reference μ records the comparison structure before value is applied. V records what the decision problem rewards. τ sets the cost of departing from the reference in units of value. Changing any one changes the decision problem.
Expected value minus τ times the departure from the reference is maximised at exactly one state. The identity behind it is an exact rewriting rather than an approximation.
The choice from reference, value, and cost is equivalent to its reference measure in both directions. A hard exclusion can remove support. Nothing in the tilt can create support the reference did not already carry.
Credal dominance is transitive and irreflexive, so a finite action set always has a maximal element. One survivor fixes the action. Several survivors are the complete decision answer until a rule, reference, or protocol refines them.
A system that returns the answer’s true shape displays greater rational power than one that compresses every problem into a point.
Intelligent Epistemology, Remark 13.1
Verdict Refusal here is lawful rather than evasive, and it comes with content: the family, and the exact boundary of what the constraints determine. A system that returns the answer's true shape is doing more work than one compressing every problem into a point. When a decision is genuinely required, credal dominance removes every dominated action and always leaves at least one standing.
▸From the paper
For a human reasoner this is suspension of judgement. For a machine it is lawful refusal. A system that returns the answer's true shape displays greater rational power than one that compresses every problem into a point.
Remark 13.1↗ · Lawful refusal carries a complete answer
The line the machine sections in Part V are built on. A model that always produces a point is not more capable than one that reports a family; it is reporting a result its constraints never bought.
Everything so far has been internal. Evidence from the world arrives through stochastic kernels: perception, memory, introspection, testimony, instruments, models. They differ in signals, errors, dependencies, incentives, and hidden source conditions. Once the kernels are specified they share one update form. A single reliability number turns out to be a special summary rather than the general case.
In plain words A signal is not evidence until you say how the world produces it. The same reading can support a claim, undercut it, or say nothing at all, depending on the instrument. Write the instrument down and everything becomes calculable. Leave it out and the number in front of you means nothing in particular.
Send a signal through and read what it is worth. The signal value is
not the evidence. What the signal is worth is a property of the kernel
that produced it. The same value arrives with three different
verdicts under three permitted kernels.
S = 1 · Λ = 4.0000 under the calibrated kernel · posterior 0.5000 from a base rate of 0.20
sendand read it under
Proposition 19.5: one signal value, three permitted kernels, three verdicts
source state ηP(s | H)P(s | ¬H)Λ(s)verdict
calibrated0.80000.20004.0000confirms
saturating0.90000.60001.5000confirms
adversarial0.20000.80000.2500disconfirms
The signal value is fixed and the three rows disagree about what it is
worth. Nothing in the token settles the
question. The answer is a property of the kernel the problem permits.
Proposition 16.3, the certificate: what the output distribution does not identify
Pr(S = 1) at (q, r)0.380000
at (1 − q, 1 − r)0.380000
signals observedlog-likelihood difference between the pair and its twin
10exactly zero
1,000exactly zero
1,000,000exactly zero
Both pairs put 0.380000 on a positive signal. They disagree by
zero to machine precision. The difference of log-likelihoods is computed
at each sample size from the counts a run of that length would produce. It does not
fall with data: it starts at zero. Self-agreement supplies
consistency. Truth calibration needs an anchor from outside the channel.
A source can be perfectly consistent with itself, report at a stable rate,
survive every internal audit, and still be inverted. The observable law is
one number and two truth models fit it exactly. This is what bootstrapping
and easy knowledge come to when the channel is written down: not a mistake
in the reasoning, but a fact about what the reasoning has to work with.
proved the channel is Definition 16.1,
the likelihood-ratio account of evidential force is Proposition 19.5, and
the non-identifiability witness is Proposition 16.3 (§16, §19.4).
argued the sensitivity,
specificity, base rate, and bias settings are display choices: the paper
states the identities and names no numbers for them. Every figure above is
computed for the kernel on screen. The log-likelihood row uses the expected
counts a run of that length would produce, and the difference it reports is
zero because the two pairs induce one Bernoulli parameter, not because the
row rounds.
▸The measured residual
the two pairs on a positive signal
5.551e-17
Each of these is one quantity computed two ways from the state on
screen, with neither route derived from the other. A figure of 0 means
the two routes landed on the same double. Anything else is the width of
one. Where a line above reports how far two readings disagree, this is
the figure it stands for.
Set sensitivity, specificity, base rate, and the source state, then send one signal through. The same signal value confirms, disconfirms, or says nothing depending on the kernel. The lower panel is the identifiability certificate: two truth-prevalence and accuracy pairs producing one observable law, with no signal able to separate them (Def. 16.1, Prop. 16.3, Prop. 19.5).
The turn
One object for six kinds of source
An evidence channel is a hypothesis space, a signal space, a space of source or apparatus states, and a stochastic kernel joining them. Observing a signal gives the posterior by Bayes' rule, subject to whatever point, family, or nonexistence structure the prior already carried. If several channel models remain possible, each is carried forward.
Perception and instruments are kernels from world states to registrations. Illusion, masking, calibration drift, and limited resolution are modelled by the same object. Memory is a temporally extended reconstruction channel whose latent state can include retention, interference, present cues, and rehearsal. Introspection is fallible in its classification even where the occurrence is immediate. Testimony moves odds by a likelihood ratio in which competence and honesty contribute differently. Adversarial sources need a kernel conditional on goals, incentives, and adaptation to the evaluator. A priori consequence stays internal throughout. It always was.
Unknown channel parameters are ordinary hypotheses. Calibration data update their joint posterior alongside the world hypotheses, and integrating over that posterior gives a predictive kernel richer than any raw sample proportion.
What the structure says
What a source cannot tell you about itself
Here is the section's hardest result. Its proof is two lines. In a binary symmetric channel with truth prevalence q and accuracy r, the observable rate of positive signals is qr + (1−q)(1−r). Many pairs give the same value. A pair and its double inversion give it exactly. Observed signals determine their marginal law. Nothing more.
Truth-conditional calibration additionally requires a decomposition into latent truths and a conditional kernel. Later-verified outcomes, independent controls, another calibrated route, or sufficiently restrictive structural assumptions can supply it. Repeated self-agreement cannot do it. Bootstrapping and easy-knowledge failures are cases of this non-identifiability. The diagnosis is structural rather than a complaint about circular reasoning.
What the structure says
Closure splits in two
If a body of grounds internally licenses a proposition and the consequence relation carries that proposition to another, the grounds internally license the second. The proof is one sentence. Closure is composition in the chosen consequence relation.
Route robustness is a different property. It does not travel automatically. A channel may reliably distinguish an ordinary hand-present case from an ordinary hand-absent one while failing to distinguish either from an adversarial simulation. The entailment is internal; the additional discrimination demand is a property of the kernel. Nothing in logical closure supplies it. Nothing could.
That single split is what Part IV uses on Gettier, on closure debates, on contextualism, and on scepticism. Two theorems, four pages apart. Between them they do most of the work of a literature.
The turn
Convergence, and its price
When the channel is specified, the truth is represented and supported, observations follow the sampling law, and false rivals are distinguishable, posterior mass concentrates on the represented truth exponentially fast. The proof bounds the marginal likelihood below on a small KL neighbourhood of the truth and bounds the far set above using Pinsker, then combines.
Every one of those four conditions is world-facing. Realisability holds when the true process is represented in the hypothesis space. Model enlargement is the only remedy for its failure. Distinguishability fails for observationally equivalent hypotheses, and if two hypotheses make identical predictions for all possible observations, no amount of evidence separates them. That is a limit of empirical inquiry rather than of the update rule.
The paper adds the finite-sample counterpart rather than leaving the asymptotic claim standing alone. Uniform convergence over a hypothesis class holds exactly when the class has bounded capacity, and probably-approximately-correct bounds convert capacity into sample sizes. Conformal methods trade the model for exchangeability and return finite-sample coverage from that single stated condition. The accounting is the same in each case: a guarantee holds inside the structure a method states, and the stated structure carries the price.
▸Technical · §17
What this block carriesThe posterior-convergence proof, bounded below and above, and the law of large numbers it rests on.
Law of large numbers, bounded case
Pr(|Z̄ₙ − μ| > ε) ≤ 2·exp(−n·ε² / (2M²))
Hoeffding's bound follows from Markov's inequality applied to the exponential moment together with the elementary sub-Gaussian estimate. The bound is summable in n, so Borel–Cantelli gives almost-sure convergence.
Choose a closed KL neighbourhood of the truth with positive prior mass. On a finite alphabet the log-likelihood ratios are uniformly bounded and continuous there, and the empirical frequencies converge, so the normalised log-likelihood converges uniformly on the neighbourhood. The marginal likelihood is eventually at least the prior mass times an exponentially small factor.
Once the empirical distribution is close to the truth, any hypothesis far in total variation is far from the empirical distribution too, and Pinsker converts that distance into a divergence gap. The far set therefore loses likelihood at an exponential rate.
Consistency requires the represented truth to lie in the support of the completed prior. A full-support reference helps; the constrained projection decides. Common finite affine cases preserve support when the reference is strictly positive, the feasible family has an interior point, and the attained exponential-family solution is interior.
Every condition in Part III earns its place by excluding a named rival. Read the table as a price list: the middle column is what you get back if you decline to pay.
4slots in a channel
The world or hypothesis space, the signal space, a space of source, apparatus, or context states, and the stochastic kernel that joins them. The source state may encode calibration, competence, incentives, selection effects, common causes, or adversarial control.
In a binary symmetric channel with truth prevalence q and accuracy r, the observable rate is qr + (1−q)(1−r). That pair and its inversion produce the same law. Self-agreement supplies consistency data; truth calibration needs an external anchor.
Realisability, prior support, stable sampling, distinguishability. Given all four, posterior mass on the far set falls exponentially and concentrates on the represented truth. Drop one and the guarantee goes with it.
Internal justification is closed under the stated consequence relation, because closure is composition. Route robustness transfers only when the route also discriminates the alternatives the entailed proposition introduces.
A single reliability number is only a special summary. Logical consequence remains an internal relation.
Intelligent Epistemology, §16
Verdict This is the boundary between rational method and empirical success. The internal rule fixes how evidence changes belief. Channel, support, realisability, and distinguishability decide whether that process finds the truth, and all four are world-facing conditions the reasoner does not control. The convergence theorem holds when they do; §18 lists what survives when they do not.
▸From the paper
Here is the boundary between rational method and empirical success. The internal rule fixes how evidence changes belief; channel, support, realisability, and distinguishability determine whether that process finds the truth.
Printed immediately after the convergence figure. It is the most load-bearing sentence in Part III. Every reconstruction of a knowledge problem in Part IV runs through it.
▸The proof idea · Thm 17.2argued
Posterior convergence
HypothesesObservations i.i.d. on a finite alphabet under a specified channel, with the prior giving positive mass to every KL-neighbourhood of the true per-observation distribution.
Fix a small closed KL neighbourhood of the truth with positive prior mass.
On a finite alphabet the log-likelihood ratios are bounded and continuous there, and empirical frequencies converge almost surely.
So the marginal likelihood is eventually at least the prior mass of that neighbourhood times an exponentially small factor.
Any hypothesis far in total variation is eventually far from the empirical distribution, and Pinsker turns that into a divergence gap.
The far set therefore loses likelihood exponentially, and choosing the neighbourhood small enough makes the ratio vanish.
For a finite identifiable class a small enough radius isolates the truth, so posterior mass on it tends to one.
The paper states 122 numbered results, remarks aside, across
33 sections. Each column below is one section. Each dot is a statement, in the order that section
states it. An arc runs from a statement back to the earlier one it cites by number or
names by title.
The paper states 122 numbered results, remarks aside, across
33 sections. Every one of them is listed below, searchable, in the
order the paper states them.
Open one and it gives its text, its proof, what it rests on, and what rests on it.
Every reflexive inferential relation whose language contains at least one consistent proposition contains a consistent inference.
The proof, as the paper gives it
Let P be a consistent proposition. Reflexivity gives the valid instance P ⊢ P. Its premise set is consistent, so it is a consistent inference and MU follows.
A theorem. A proposition and a corollary are the same mark, smaller and lighter.
An open ring is a lemma. The smallest ring is a definition.
A ringed dot is one of the 7 landmarks named above.
A citation: the paper prints the number. 27 of them.
A name: a statement or its proof uses, word for word, the title the paper gave an
earlier one. 25 of them.
Nothing else is drawn. Prose that gestures at a result without naming or numbering it
is left alone. A guess in a dependency graph is a false edge.
52 statements stand in at least one of the two relations. The other
70 are drawn quiet. A trace ends where the numbered references stop. The paper
declares no axioms; that is the floor this map can show.
The 35 remarks are not here. §1.1 grades them as locating rather than
proving. Neither are the 49 argument blocks, which carry no number
and are cited by subsection.
Source: Intelligent Epistemology — MU and Epistemic Zero. Statements and proofs are the
paper's own words, read out of the LaTeX source. The layout is a layered DAG, computed once at build.
Part IV
Classical Problems: Resolutions, Reconstructions, and Boundaries
Part IV
The register changes here. Parts I to III prove and define; from §19 the paper argues, and forty-nine of its argument blocks live in this part and the two after it.
Our knowledge can only be finite, while our ignorance must necessarily be infinite.
The Münchhausen Trilemma (Albert, 1968) classifies inferential support as regress, circle, or unsupported stopping. MU opens a fourth: a finite direct existence proof that terminates. From that foothold the section works through Carroll's tortoise, rule-following, the myth of the given, indifference, the lottery, Dutch books, Moore's paradox, and the structure of knowledge itself.
In plain words The old trap says every justification either goes on forever, curls back on itself, or stops somewhere arbitrary. There is a fourth option nobody used: a short proof that finishes. Once you have that, a lot of famous puzzles turn out to be asking two questions at once. They come apart cleanly.
The classification offers three ways for support to run out: it goes
on for ever, it comes back round, or it stops on an assertion. Walk
each one and read the ledger. Then walk the fourth, which the
classification never enumerated.
MU ≡ ∃I CI(I)P ⊢ P, then ∃-introduction
regress · 0 moves · 1 support demand open · 0 discharged
moves taken0steps the reader has walked
demands open1support still owed
discharged0demands closed, not relocated
established0propositions beyond the premises
accepted unproved0stipulations in the walk
hypotheses in force1a support relation on propositions
The two right-hand cells measure different things. A stipulation is
something the walk accepts as it goes. A hypothesis is what the case stands
on before it moves: the fourth case discharges its support demand and still
runs on Theorem 2.1's two.
One claim, one open support demand. Take a step and watch where the demand goes.
The denial, separately: assert ¬MU, that no consistent inference exists (Cor 2.2)
attempts: 0consistent denials found: 0
Suppose ¬MU is consistent. It is then a consistent proposition of the language.
Reflexivity applies to it like any other. ¬MU ⊢ ¬MU is a valid instance.
Its premise set is consistent. The instance is a consistent inference.
That inference witnesses ∃I CI(I), which is the existence ¬MU denies.
Not attempted. The steps above are the whole argument; press the button
and the card walks them with the counter running.
Each step discharges nothing. The support demand moves to a fresh proposition and arrives there intact, so the count of open demands is one after every move ever made. Nothing about the walker fails. The chain has no end to reach.
proved the classification and the
fourth case are Proposition 19.1; the fourth case's three moves are the
proof of Theorem 2.1, and the denial panel is Corollary 2.2 (§19.1, §2).
The fourth case's two hypotheses are Theorem 2.1's own, taken from its
statement.
argued counting one hypothesis
against each of the three horns reads the classification as a
classification of inferential support, which is how §19.1 states it. The lane drawing, the
three-node loop, and the three-step stopping walk are display lengths
chosen to fit the frame. The counts they feed are computed from the walk itself. The regress
carries a named display clamp at 40
moves — the walker stops there and the chain does not.
located Corollary 2.2 holds in any
consequence structure satisfying Theorem 2.1's hypotheses whose language
contains the proposition ¬MU. A paraconsistent consequence relation
defines a different theorem-generation problem with its own closure
(Remark 3.3). The denial is not refuted there; it is relocated.
Walk each horn and read its ledger, then walk the fourth. Proposition 19.1's claim is exact and narrow: within the consequence structures the MU Theorem covers, the grounding of MU is none of the three. One consistent proposition, one application of reflexivity, one existential introduction (Prop. 19.1, Thm 2.1, Cor 2.2).
The turn
The fourth case
Regress, circle, dogma. Three horns. The trilemma has organised the theory of justification for half a century by insisting those are the options. Proposition 19.1 states the escape precisely: within the consequence structures the MU Theorem covers, the grounding of MU is none of the three. It is a direct existence proof.
The proof has one reflexive inference instance and concludes existentially that at least one consistent inference exists. It terminates. That is the escape. Its premise is an instance of reflexivity, and the chosen proposition disappears under existential introduction. The reflexive analysis of inference begins only after that result is in hand. That order is what keeps the argument non-circular.
The problem of the criterion (Chisholm, 1973) is handled the same way, by role separation rather than by a stronger argument. Method comes from what it is for an answer to come from grounds. Premises, models, representations, and channels supply the material to which the method applies. The circle dissolves and the empirical work of improving the material stays exactly where it was.
What the structure says
Carroll's tortoise, and the rule that will not become a premise
The tortoise accepts a valid argument's premises and demands one further premise before accepting the conclusion. Grant the demand and the same demand reappears immediately, since deriving the conclusion from the enlarged premise set still requires applying a rule. Adding a premise about that application reproduces the role at the next level. The construction iterates forever. No premise ever becomes a rule.
The regress arises from asking a premise to perform the role of the consequence relation. Rules and premises occupy their proper places, and a consequence relation can still be challenged, compared, or revised inside another inferential setting whose grounds and rules are explicit. That closes the regress at the level of rule use.
The other rule-following problem concerns continuation rather than application. A finite record on an infinite domain, an unobserved case, and at least two available values are enough to build two total functions agreeing on everything observed and differing where nothing was. The proposition settles the exact underdetermination present in the record. A unique continuation then needs further structure: a representation, a simplicity or invariance standard, a communal practice, an explicit algorithm.
What the structure says
Experience, and what a signal is worth
Experience is often asked to be an input that supports belief and to carry its own complete interpretation. The first role is available. The second is too strong. The proof of that is a page of arithmetic.
Take an experiential signal and a hypothesis. The evidential force of the signal is the likelihood ratio under the channel and background. Pick two conditional probabilities and the ratio can be made greater than one, less than one, or exactly one. The same sensory token can therefore support the hypothesis, oppose it, or say nothing, and the token alone fixes none of these.
The result locates immediate defeasible support (Pryor, 2000) in the standing relation between signal, background, and channel. An experience enters the problem as a present signal and shifts belief through a standing perceptual channel. Its classification and the channel remain revisable. First-order support and second-order certification are different questions on different levels.
The turn
Indifference returns a shape, not a number
The classical indifference paradoxes are the four shapes of Definition 3.4 in disguise. Take a cube whose side length lies between zero and one. Uniformity in side gives one probability, uniformity in face area gives another, uniformity in volume gives a third. Each is coherent relative to its own description. The three answers reveal a reference family, and the visible constraints preserve all three.
Nor is this special to cubes. Bertrand's chord problem has the same structure, since distinct physical procedures for drawing a random chord induce distinct measures. A physically symmetric die is the contrasting case. Its permutation symmetry among six faces is carried by the problem. The uniform answer is fixed directly. Symmetry carried by the problem itself determines a point. Symmetry that has to be supplied by a choice of description determines a family.
The classification has four outcomes and the paper names all four: a unique reference and hence a point; a family of reference-relative answers; a nonattainment result; or no admissible case at all. The last three become point-valued only when an enriched problem supplies the missing reference, attainment, or consistency structure.
What the structure says
Moore's paradox, and the level the oddity lives on
Moore's paradox lives on the assertion level, not on the proposition. The conjunction of p with the agent's not believing p can be true: it may be raining while an agent fails to believe that it is raining. Suppose a sincere assertion of p is evidence that the speaker believes p. Then a sincere assertion of the conjunction asks the assertion channel to represent the speaker as believing p and as denying that belief at once. The proposition can be logically consistent while its sincere assertion remains unstable as a truthful self-report.
The neighbouring sentence, p but my evidence does not support p, can be coherent when the first clause reports a truth learned through another route. It becomes epistemic akrasia when the speaker treats the second clause as an authoritative assessment of all current grounds and goes on endorsing the first without further reason. The conflict is represented as disagreement between two channels or levels of one model. Treat the higher-order judgement as certain and authoritative, and retaining the conflicting first-order state violates the retention law. Leave it uncertain and the output can remain a joint distribution over the claim and over the reliability of both assessments.
What the structure says
Thresholds, bridges, and the profile
The lottery separates high probability from deductive certainty and from knowledge. A ticket is overwhelmingly likely to lose, and its loss stays graded belief. For every threshold below one there are coherent cases where the probability exceeds it while the background does not entail the claim. The preface runs the other way: an author can rationally be confident in each sentence and expect that at least one is wrong, because conjunction aggregates the small risks its members carry.
Three bridges often conflated with the probability calculus get their own conditional treatments. Dutch-book coherence follows from a fair-price protocol with unrestricted combination and linear valuation. The chance–credence bridge follows from calibration plus admissibility. That pair is the conditional core of the Principal Principle. Reflection follows from the martingale property of a posterior in one model, and it breaks under anticipated forgetting, model change, strategic distortion, or future irrationality.
The section ends by replacing the search for a scalar essence of knowledge with a profile. Truth, internal support, and route integrity are logically independent, so the classical cases occupy different vertices: a lucky guess, a demon world, a Gettier case, a reliable clairvoyant, and ordinary knowledge at the far corner. No analysis based on one coordinate captures the contrasts that motivate the cases, which is the whole argument of a fifty-year literature stated as a cube.
3horns, and a fourth case
The Münchhausen Trilemma classifies inferential support as regress, circle, or unsupported stopping. MU opens the prior fourth case: a finite direct existence proof that terminates by existential introduction.
Add the premise that A and A → B license B. Deriving B from the enlarged set still uses a rule. A further premise reproduces the role at the next level, and the construction iterates without end.
A cube with side length in [0,1]. Uniformity in side gives one half. Uniformity in face area gives a quarter. Uniformity in volume gives an eighth. Each is coherent relative to its description, and the visible constraints preserve all three.
Truth, internal support from the stated grounds, and the reliability or safety of the route across the relevant class of cases. The three vary independently, so no analysis based on one of them captures the classical contrasts.
The proof terminates. Its premise is an instance of reflexivity from reflexivity and a consistent premise.
Intelligent Epistemology, Proposition 19.1
Verdict Note what is not claimed. MU's proof escapes the trilemma; individual empirical beliefs still need their grounds, their models, and their channels. What the fourth case buys is a floor to stand on while assessing them. The assessment is itself an inference governed by the relation MU concerns. The paper says so explicitly and treats it as a feature.
▸From the paper
The Münchhausen Trilemma classifies inferential support as regress, circle, or unsupported stopping (Albert, 1968). MU opens the prior fourth case: a finite direct existence proof.
§19.1↗ · Regress, the Trilemma, and the Problem of the Criterion
Prior is the load-bearing word. The fourth case is not a fourth way of justifying a belief; it is an existence proof that runs before the analysis of justification begins.
▸The proof idea · Prop 19.1argued
Grounding result
HypothesesThe consequence structures covered by the MU Theorem: a reflexive inferential relation whose language contains at least one consistent proposition.
The proof of MU has one reflexive inference instance, P ⊢ P, for a consistent P.
It concludes existentially that at least one consistent inference exists, so the proof terminates.
Its premise is an instance of reflexivity, taken from reflexivity and a consistent premise, so nothing is assumed unproved at the halt.
The chosen P disappears under existential introduction, so the conclusion does not derive MU from MU.
The reflexive analysis of inference begins only after that result is in hand, which is what keeps the order non-circular.
Induction, the ravens, old evidence, Goodman's riddle, abduction, Gettier, testimony, disagreement, scepticism. Nine famous problems, and one distinction running under all of them: internal support and connection to truth answer different questions. Draw them apart and each problem states its own remaining work rather than resisting solution.
In plain words Being right and reasoning well are two different achievements. You can do one without the other. Most of the classic puzzles about knowledge are built on that gap. Once you stop asking a single notion to cover both, the puzzles turn into ordinary questions about evidence and about how reliable your route to the fact was.
The knowledge profile, after the paper's own plate. Truth, internal support, and route integrity are independent axes, so the classical cases sit at different vertices: lucky guess, demon world, Gettier, reliable clairvoyant, and knowledge at the far corner.
The turn
Induction keeps its standpoint
The classical demand asks for a justification of inductive practice from outside all inference. The reply is short and it is structural. Every reason offered for accepting or rejecting a projective rule is itself an inference governed by a relation of support, so every proposed tribunal is an inference too. Assessment stays reflexively inside the practice. There is no outside.
That settles the constitutive half and hands the rest to empirical work. The future-facing result follows from the projective structure, support, sampling, and distinguishability stated in the problem. Realisability, stability, support, sampling, and distinguishability carry the burden. Where realisability fails, updating can compare only the candidates present in the model class. At best it concentrates on a predictively optimal region under further misspecified-learning conditions.
The paper credits Strawson with the constitutive half of the diagnosis, and then names what his reassurance could not deliver: the convergence conditions. Asking whether induction is rational does resemble asking whether the law is legal. The resemblance does not tell you when a particular projective rule will keep working.
What the structure says
Two paradoxes, one missing model
The ravens paradox arises from combining a logical equivalence with an unstated confirmation model. Evidence confirms a hypothesis exactly when it is more probable under it than under its negation. The likelihood ratio is fixed by the sampling and population model. A nonblack nonraven can therefore confirm, disconfirm, or leave the hypothesis unchanged across different models. Often it confirms minutely, because ravens are rare. The direction is still real.
Sharper still, confirmation follows the hypothesis specified. The paper builds an explicit four-cell model in which evidence confirms a hypothesis while disconfirming a conjunction containing it, which is enough to show that logical containment does not transmit confirmation.
Old evidence is the mirror case. If the agent's current state already assigns probability one to the evidence, conditioning on it does nothing. What is learned when a new theory gains support is the theory-to-evidence relation. The odds move on that relation's likelihood ratio, while the old observation stays inside the current state where it already was.
What the structure says
Representation before projection
The green–grue construction (Goodman, 1955) shows that data support projection only through a represented language. Complexity is representation-relative: a predicate simple in one language may be elaborate in another. A projective inference becomes well defined through its representation language, measurement apparatus, invariances, and candidate transformations. Several permitted representations produce their complete family.
No Free Lunch supplies the computational counterpart in its own finite setting: uniform averaging over an unrestricted problem class equalises the performance of every learner. Successful generalisation therefore requires problem-relative structure. Successful induction makes its bias explicit as exactly that structure.
At deployment scale the riddle recurs as distribution shift. A rule fitted to cases examined before some boundary gets projected onto cases beyond it. The choice among extrapolations lives in representation and apparatus rather than in the record. That sentence is the bridge from a 1955 puzzle to a live engineering problem.
The turn
Gettier, and the third coordinate
Gettier cases show that justified true belief can be true through the wrong connection. The failure is structural. Internal support and connection to truth answer different questions. A case can score on both while the connection between them is accidental. Adding a route condition closes the gap.
A belief-forming route is robust relative to a chosen world class when its kernel keeps discriminating the truth-relevant alternatives throughout that class to the standard the application requires. The choice of class and standard belong to the problem. That is why contextualist and relevant-alternatives insights fit here without a semantic thesis. A route may count as knowledge-producing relative to one class and standard and fail relative to a more demanding pair. The attributions stay consistent because the problems differ.
Two further pressure cases fall out. A subject in a perfectly deceptive demon world can be internally indistinguishable from a normally situated counterpart while lacking the external success. A reliable clairvoyant (BonJour, 1980) with no accessible reason to trust the faculty has the external route without the internal support. Neither coordinate replaces the other. Both are needed.
What the structure says
Testimony, disagreement, and where doubt lands
Bootstrapping cases use a source's outputs to assess that source's reliability, and the channel model names the missing ingredient exactly: an external anchor. Signals identify truth-reliability only through independently verified outcomes or sufficiently restrictive structural assumptions. Repeated self-agreement supplies consistency data. It stops there.
Testimony earns its weight through a Bayes factor in which competence, honesty, dependence, incentives, and selection are separate features. Expert testimony can be rationally weighty without the hearer reproducing the proof, since division of cognitive labour is itself a calibrated channel structure. Dependence matters more than volume. Ten reports copied from one source form one likelihood structure.
Peer disagreement is evidence about another route and often about one's own. An independent, comparably reliable peer supplies a substantial factor. A report determined by the same evidence and method supplies a factor of one and reveals a disagreement to be explained. Conciliation and steadfastness follow from the peer channel, the shared evidence, and the dependence structure in each case rather than from a general policy.
Global scepticism about inference is defeated by the MU Theorem. What remains is local and tractable: Cartesian, brain-in-a-vat, perceptual, testimonial, and scientific doubts become coherent questions about a channel, a model class, or a connection to the external world. A perceptual route may robustly support ordinary claims across ordinary nearby cases while failing to distinguish them from perfectly adversarial simulations. The larger anti-sceptical claim asks more of the route.
4projectibility conditions
Empirical fit under the channel model. Stability under the chosen transformations. Complexity relative to the chosen coding or reference structure. Predictive performance on held-out or future observations. All four are relative to a fixed representation.
In the paper's worked model, evidence raises the odds on the hypothesis because 0.145 exceeds 0.01. The same evidence lowers the posterior of the conjunction from 0.45 to about 0.290. Confirmation follows the specified hypothesis rather than the logical equivalence.
Two model bundles with equal signal laws under every permitted history and action. The product of ratios is one, so no finite history from any adaptive policy moves their prior odds.
True; internally supported by the problem; route robust in the chosen relevant-alternatives class; the success attributable to the relevant competence or channel. This is offered as a proposal for protection against epistemic luck.
Every proposed tribunal is itself an inference, so inferential assessment remains reflexively inside the practice.
Intelligent Epistemology, §21.1.1
Verdict The pattern repeats with variations. Induction keeps its standpoint and hands the projective task to explicit models. The confirmation paradoxes turn on unstated sampling models. Goodman moves to representation. Gettier gets a third coordinate. Scepticism splits into a self-defeating global version and a tractable local one. In each case what remains is stated rather than dissolved.
▸From the paper
Gettier cases (Gettier, 1963) show that justified true belief can be true through the wrong connection. A right answer can still be bad inference. The failure is structural: internal support and connection to truth answer different questions.
Three sentences and the middle one is the shortest in the paper. The structural diagnosis is what makes the knowledge profile inevitable rather than stipulated.
The section gathers the results in one table. Each row names a classical problem and states what follows, in the paper's own verdict vocabulary: resolved, reconstructed, located internally, reconciled, substantially explained, proved within an experiment class, bounded by the empirical quotient, or separated. Then it classifies the residue into five forms, so what remains has a shape rather than a shrug.
In plain words Here is the scoreboard. Some problems are settled outright. Some are rebuilt so they can be worked on. Some are shown to have a boundary that no amount of thinking will cross. The last part is the useful bit. Whatever is left over falls into five kinds. Each kind tells you what to do next.
Direct resolutions establish inference as its own standpoint, rule application as irreducible to an added premise, thresholds as context-indexed, and channel calibration as externally anchored. Each is a claim that something classical was asking the wrong question. Each comes with the proof that shows why.
Structural relocations move a problem somewhere it can be worked on. Induction goes to projectibility and sampling. Goodman goes to representation. Local scepticism goes to models and channels. Disagreement goes to evidence, dependence, and reliability. Nothing is dissolved by relocation. The problem acquires an address.
Positive reconstructions supply accounts where there was a gap: abduction, empirical knowledge, testimony, scientific inquiry, and the norms internal to inference. These are the rows where the paper builds rather than clears.
What the structure says
Five forms of remaining work
The residue is classified rather than gestured at. Candidate generation: inquiry enlarges the problem with a representation, hypothesis, or rule. Empirical discrimination: new channels or experiments separate the surviving cases. Computational access: further computation reaches more of a fixed full answer. Practical value: value, loss, cost, or an ambiguity rule selects among surviving options. Empirical or metaphysical equivalence: inquiry records the class until a new discriminating structure appears.
Each form has a matching answer type and a matching earlier theorem. A higher-order family. A probability or credal state on an empirical quotient. An approximation contract. An undominated decision set or value-relative completion. A proof of non-identifiability. The classification is useful only because the theorems are already there to receive it.
What the map claims is modest and unusual. It does not say the classical problems have been solved. It says less than that. It says each one has been pushed until it states its own remaining work. The remaining work has five shapes rather than an unbounded number. Five shapes. Generate a candidate. Run a discriminating experiment. Compute further. State the value at issue. Mark the boundary. Each proceeds under the same rule of inference that got the problem here.
Eighteen rows from the map, in the paper's own verdicts (§23.1)
Eighteen rows from the map, in the paper's own verdicts (§23.1)
Problem
What follows
Grounding regress and the Münchhausen Trilemma
Resolved at the level claimed. MU has a finite direct proof that terminates by existential introduction and precedes the later analysis of inference.
Problem of the criterion
Resolved by role separation. The inferential form comes from what it is for an answer to follow from grounds; premises, representations, and channels supply the material.
Rule-following and Carroll's regress
Resolved in two layers. Rule application is irreducible to an added premise, and finite behaviour supports a family of continuations.
The Given and immediate experience
Reconstructed. Experience supplies immediate defeasible support through a standing channel and model that connect signal to claim.
A priori support and analyticity
Located internally. Meanings, axioms, and rules determine a priori support within a formal problem; empirical inquiry assesses the framework's relation to the world.
Indifference, reference classes, and self-location
Resolved by answer shape. Real symmetry determines a point. Several measures, reference classes, or observation protocols produce their full family.
Uniqueness and permissivism
Reconciled. One fully defined problem has one full answer, and that answer may contain several permissible credal states.
Reconstructed through self-representation. The Moorean proposition may be true while its sincere assertion conflicts with the speaker's represented belief.
Induction
Resolved in two stages. Inference supplies its own standpoint. Projective rules are then compared through explicit models, support, sampling, and distinguishability.
Goodman's new riddle and No Free Lunch
Resolved at the epistemic level. Representation and problem distribution supply the projective structure; evidence compares the cases that make different predictions.
Gettier, closure, context, and stakes
Resolved through the knowledge profile. Internal support and route integrity are independent coordinates that jointly support empirical knowledge.
Global and local scepticism
Resolved by scope. MU settles the possibility of inference. Doubt about a channel, model, or world-connection becomes a coherent local problem.
Meno and the swamping problem
Substantially explained. Information has nonnegative expected instrumental value, and knowledge adds a stable, reusable, attributable route to truth beyond lucky arrival.
Duhem–Quine underdetermination
Proved within an experiment class. Equal signal laws preserve posterior odds across every available test. New interventions or assumptions enlarge the discriminating class.
Scientific realism, pessimistic meta-induction, and unconceived alternatives
Bounded by the empirical quotient. Data identify models up to observational equivalence; model enlargement introduces unconceived alternatives.
Logical omniscience and bounded rationality
Separated. Deductive closure is the full normative answer, while a finite agent reaches a computably accessible subset at each time.
Machine understanding and opacity
Resolved at the epistemic level and bounded at the phenomenal level. Reliability, route integrity, and auditability determine epistemic standing.
Eighteen of thirty-one, in the paper's order, picked so that every verdict in the vocabulary appears at least once. Ten of the eighteen open with some form of Resolved; in the full table twenty of the thirty-one do. The other seven verdicts are Reconstructed, Located internally, Reconciled, Substantially explained, Proved within an experiment class, Bounded by the empirical quotient, and Separated. The full table runs from page 60 to page 63, and those eight verdicts are eight different claims.
31classical problems, tabled
From the grounding regress to machine understanding. Each row names the problem and states exactly what follows, in the paper's own verdict vocabulary rather than a summary written for it.
Direct resolutions establish inference as its own standpoint, rule application as irreducible, thresholds as context-indexed, and channel calibration as externally anchored. Structural relocations move a problem into representation, sampling, or dependence. Positive reconstructions supply accounts.
Candidate generation, empirical discrimination, computational access, practical value, and empirical or metaphysical equivalence. The corresponding answers are a higher-order family, a state on an empirical quotient, an approximation contract, an undominated decision set, and a proof of non-identifiability.
The scientific, creative, practical, and metaphysical work remains to be done.
Intelligent Epistemology, §23.2
Verdict The five forms are the paper's own residue classification, not a list of open problems written afterwards. Each has a matching answer type and a matching theorem: candidate generation goes to the higher-order problem, empirical discrimination to the quotient theorem, bounded access to the computational analysis, practical underdetermination to the decision results, and empirical equivalence to the channel model.
▸From the paper
The gain is knowing what kind of work it is: generate a candidate, run a discriminating experiment, compute further, state the value at issue, or mark the boundary. Each proceeds under the same rule of inference.
The closing line of Part IV, and the paper's own answer to the charge that a theory of principle explains nothing. It does not finish the work. It tells you which of five things the work is.
Part V
Science, Computation, and Society
Part V
Six sections of the paper, §24–§29, in one here: falsification, realism and the empirical quotient, bounded rationality, social epistemology, artificial intelligence and causal models, and the scientific cycle.
We can only see a short distance ahead, but we can see plenty there that needs to be done.
Experiments, finite computation, social dependence, and artificial systems. The same distinctions now meet the places where knowledge gets made. Falsification becomes a region of one likelihood comparison. Realism reaches the quotient its experiments fix. Alignment turns out to be reference-based as well as value-relative. And a generative system that reports content its constraints never paid for gets a name.
In plain words This is the part where the machinery meets practice. What a severe test is. How far the evidence lets you go about what is out there. What a bounded reasoner owes when it says it is approximating something. Why ten reports copied from one source are one report. And what exactly is wrong when a model states a fact it never had grounds for.
A test is severe for two alternatives to the extent that the resulting signal laws are distinguishable and the apparatus state is controlled. A failed severe prediction can strongly reduce support. A passed one increases support in proportion to how much less expected it was under the rivals. Falsification and confirmation are complementary regions of one likelihood comparison. They are not two epistemologies.
A failed prediction confronts a bundle: theory, auxiliaries, apparatus, and background conditions. Revision localises the change the signal supports and records every alteration of the problem. Popper's insight survives as severe exposure. Duhem's correction is absorbed into a fully specified test package. Kuhn's complication becomes an explicit translation problem, since rival theories may organise salience, measurement, and even the space of candidate questions differently.
What the structure says
How far realism reaches
The observational-equivalence theorem applies directly. Take a class of possibly adaptive experimental policies and call two models equivalent when they induce the same law on complete observation histories under every policy. If both have positive prior probability, every history with positive shared likelihood leaves their posterior odds exactly where they started. Evidence from that class identifies models at most up to the quotient.
The corollary sharpens it. Every property learned from data generated by the class is constant on each equivalence class. Empirically decidable properties are exactly the class-invariant ones. The maximal empirical answer is a probability or credal state on the quotient, together with the structure invariant inside its classes. Structural realism gets a precise statement out of this: where changing theories preserve structure across domains of successful prediction, and that structure is class-invariant, it is the strongest realist content those data support.
Unconceived alternatives get a proof rather than a worry. Here it is. An alternative outside the considered class receives no posterior probability. In an enlarged class, assigning zero prior mass keeps it at zero after every finite update, because Bayes multiplies and renormalises. High posterior concentration on one member therefore establishes comparative success inside the class, and nothing more, unless realisability or class adequacy is separately supported.
What the structure says
What a bounded reasoner owes
An exact theory gives bounded reasoning a target. When a bounded procedure is presented as approximating an exact inferential answer, the claim has to state four things: the exact problem and target answer, the resource budget, an error or regret or calibration or convergence guarantee, and the conditions under which the guarantee holds. A practical procedure may instead be evaluated directly by a loss, a calibration test, or a benchmark. Either way there is a contract. State it.
Deductive closure is the normative answer; the consequences a bounded agent has reached by some time are a subset of it. Further valid computation enlarges the subset without changing the target. That separation makes learning a proof, or a theory-to-evidence relation, genuinely new information for the agent even when it was implicit in the ideal closure.
The paper then applies the contract to a list without exempting anything. A variational approximation states its family, objective, and bound. A Monte Carlo method states its target measure, transition kernel, and mixing control. A heuristic search states its objective, moves, stopping rule, and failure modes. The free-energy principle is treated the same way: a model family with an objective, whose epistemic standing is priced by the channel, reference, and bound it states.
The turn
Dependence, credibility, and what a marker predicts
Agents with a common prior whose posteriors are common knowledge cannot agree to disagree. The scope of that theorem is exactly those two conditions. Absent them, rational disagreement can locate a difference in evidence, reference families, model classes, source dependence, or computational access. Differences in utilities produce disagreement about what to do while the credences agree.
Reports that are conditionally independent given the truth and their source states have a factorising joint likelihood. Reports deriving from one common source, dataset, or coordinated incentive do not. Multiplying them as independent double-counts the evidence. Rational deference therefore depends on competence, independence, conflict of interest, and the hearer's access to calibration data.
The paper states one component of testimonial injustice as a dependence condition. If the truth of a reported claim is independent of a social marker given the source-relevant evidence, then changing a source's weight solely because of that marker makes the answer depend on information with no predictive relevance. Where the marker is predictive only because social structures affect access or treatment, the predictive fact and the moral evaluation of the structure are kept distinct. The formalism represents credibility deficits, testimonial suppression, unequal access, and institutional dependence. Social and moral premises determine their evaluation and remedy.
What the structure says
Machines: prediction, policy, and the standard
Universal induction is indexed to a reference machine. The paper does not soften the point. Once the coding machine is fixed a powerful result follows. Changing the machine changes the case, and adversarial choices can produce pathological universal priors. Leike and Hutter's bad-prior constructions are cited as a decisive control on any finite-scale claim of universality.
AIXI is decomposed into four logically distinct inputs: the algorithmic environment semimeasure, the reward channel, a horizon or discount convention, and expectimax optimisation. Universal induction constrains the first relative to a machine. The other three are separately specified. Prediction and policy stay apart.
Alignment gets the same treatment through the choice theorem. At a fixed decision state the KL-regularised policy is the reference policy reweighted by the exponential of action value, and the components stay distinct: reference support sets the available action class, reward ranks that class, the cost parameter governs departure from the reference, and the optimisation protocol determines implementation. A misspecified or truncated reference can exclude relevant actions. Reward misspecification and reward hacking live in the second slot. Both are old problems in a new place. Too small a cost permits over-optimisation and too large a cost leaves the reference dominant.
Then the standard, applied without exception. A generative system that reports content its constraints never paid for instantiates smuggling at scale. Non-Smuggling applies to machine reporters unchanged. No exemption is available. Hallucination is its violation under thin constraints. Accuracy and auditability remain independent dimensions, so a system may be accurate enough to deserve weight while being too opaque for high-stakes delegation, because action also requires governance, contestability, and responsibility.
▸Technical · §25 · §28
What this block carriesThe two calculations that carry Part V: the likelihood cancellation behind the empirical quotient, and the bidirectional Gaussian regression.
The quotient cancellation
Empirical equivalence makes the history likelihoods equal under every allowed policy. The policy contributes the same action-selection probability under both models because it is a function of the shared observed history. Bayes leaves the posterior odds at the prior odds. Adaptivity changes the policy without touching the equality.
Regress one jointly Gaussian variable on the other and the residual is independent of the regressor. Regress the other way and the same is true. Both structural equations generate the identical joint observational law, so nothing in the distribution orients the arrow.
At a fixed decision state the choice theorem gives the KL-regularised policy directly. In a sequential soft-control problem the formula applies statewise once the soft Bellman recursion has determined the action value. The one-step tilt is a component of the sequential construction rather than the whole of it.
Transcribed from the paper, which marks every one of these proofs
optional on a first reading — and prints them anyway.
The full text ↗
ℳ / ≡_𝒜what the evidence identifies
Models are identified at most up to empirical equivalence within the available class of experimental policies. Empirically decidable properties are exactly the class-invariant properties.
An alternative outside the considered model class receives no posterior probability. In an enlarged class, zero prior mass survives every finite Bayesian update. That is the force of the problem of unconceived alternatives.
Every nondegenerate correlated bivariate Gaussian admits linear structural representations in both directions with independent disturbances. Causal direction therefore requires identifying structure beyond the observational distribution.
A generative system that reports content its constraints never paid for instantiates the failure Non-Smuggling names. The standard applies to machine reporters unchanged. Hallucination is its violation under thin constraints.
A certified boundary is a finding, and a lawful refusal is an answer with its reasons attached.
Intelligent Epistemology, Conclusion
Verdict Two boundaries organise the whole part. Evidence within an experiment class identifies models only up to empirical equivalence, so class-invariant structure is the maximal empirical content. And zero prior mass stays zero under every finite update. A newly conceived alternative gains standing only after the problem is enlarged. Both are proved. Both apply to human and machine inquiry without modification.
▸From the paper
A generative system that reports content its constraints never paid for instantiates smuggling at scale; the Non-Smuggling standard applies to machine reporters unchanged, and the failure called hallucination is its violation under thin constraints.
One sentence connects Corollary 3.5 on page 6 to the most-discussed failure mode in contemporary machine learning. The diagnosis is not that the model is unreliable; it is that the output claims support the constraints never delivered.
▸What this paper grounds
Within the finite and measurable domains stated in Intelligent Epistemology, dependence on grounds generates probability and exact informational accounting generates relative entropy. A permitted reference and constraints generate least-informative completion; a current state and new constraints generate retention by KL projection.
Intelligent Systems — The Adaptive Closure · Inherited Result 4.1
The descendant states its debt in its own words. Intelligent Systems imports this paper's four laws as an inherited result, adds a value functional and a departure price, and then asks what survives when the resulting operator is composed with itself, with other systems, with observation, with scale, and with time.
Part VI
Norms and the Value of Knowledge
Part VI
If any man is able to convince me that I do not think or act right, I will gladly change; for I seek the truth, by which no man was ever injured.
Present an output as inference and you are bound, in that capacity, by explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. No moral premise is required to get there. Whether to enter an inquiry at all, how much effort to spend, and how truth ranks among other goods are separate questions with separate premises.
In plain words If you say something follows from your reasons, you have already accepted the rules for that kind of saying — the way calling something a proof accepts the rules of proof. That is not a moral demand imposed from outside. It is what the claim already meant. Whether you should have looked into the question at all is a different matter.
Three kinds of claim, kept apart. Descriptive says how agents in fact reason. Constitutive says what must hold for an activity to be that activity. Practical supplies a reason to undertake it. The norms of inference sit in the middle band and need no premise from the third.
The turn
Reasons to act, and reasons to believe
Two agents share a belief state and face different utilities. Their optimal actions may differ while every credence stays the same. A wager, threat, reward, or high stake can therefore change rational action without changing evidential support. The proof takes one line. The hygiene it enforces is worth keeping.
The same distinction clarifies doxastic voluntarism. Agents choose investigations, reports, and actions; the grounds determine support. Ordinary use of the word knowledge may stay sensitive to pragmatic context. The line between support and stakes does not move with it.
Decision paradoxes are handled the same way. Newcomb's problem requires the relation among prediction, causation, policy, and utility to be stated, and evidential and causal decision theories formalise different dependence structures. The supplied structure determines which problem is being solved. That is not a dodge. It is the diagnosis.
What the structure says
Constitutive, not moral
Theorem 31.1 makes the claim precise. An agent or process that presents an output as inference is bound, in that capacity, by the standards of Part I: explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. An output also presented as complete must preserve the whole supported answer.
The proof runs through the method theorem. An output satisfies the inferential description exactly when every dependence it carries is carried by its stated grounds. The standards are therefore internal to the claims being made, in the way the rules of proof bind a purported proof and preservation laws bind a purported homomorphism.
A reasoned rejection of every inferential norm is itself offered as an inference. The epistemic standpoint is inescapable for reasoned evaluation. What that inescapability does not settle is whether to enter an inquiry, how much effort to spend, or how epistemic goods trade against non-epistemic ones. Those need an independent premise. The paper says so plainly rather than reaching for one. It does not overclaim at the end.
What the structure says
What knowledge is worth
Free optional information never hurts. After observing any signal, the best action available is at least as good as retaining the action already chosen. Taking expectations gives the inequality directly. With a cost attached, inquiry is strictly preferred when the value of the information exceeds it, rejected when it falls short, and indifferent at equality. The result is exact for the supplied utility and cost.
The Meno problem asks why knowledge is worth more than mere true belief, since both can guide the same immediate action. The route-integrity answer is that knowledge is stable under perturbation, reusable across nearby problems, and attributable to a method. A true answer reached accidentally may guide one action. Once. A truth-connected method supports navigation, transfer, and correction.
The swamping problem asks what a reliable route adds once truth is already present. Truth is the successful outcome; a robust route explains why the success recurs, survives relevant counterfactual changes, and can be credited to the agent or the channel. That supplies the extra epistemic value. Final moral value stays in the practical domain where it belongs.
Internalism and externalism close the part by answering different questions. It is the pattern of the whole paper compressed into two pages. Internal support asks whether the state is a valid answer to the agent's grounds. External success asks whether representation and channels connect those constraints to the world in the required way. A demon-world subject may be internally impeccable and externally unsuccessful; a lucky guesser may be externally correct and internally unsupported. Reliabilist, safety, competence, and virtue accounts supplement the internal analysis rather than replacing it.
3kinds of claim
A descriptive claim says how agents in fact reason. A constitutive claim states what must be true for an activity to count as that activity. A practical or moral norm supplies a reason to undertake or prioritise it.
Explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem. An output also presented as complete must preserve the whole supported answer.
After any signal the inner maximum is at least the conditional expected utility of the action already chosen. Taking expectations gives the inequality. Information that is free and optional has nonnegative optimal expected utility.
Strictly preferred above the cost, rejected below it, indifferent at equality. The result is exact for the supplied utility, cost, and option to ignore the signal. Categorical moral duties belong to the wider value problem.
A reasoned rejection of every inferential norm is itself offered as an inference and is therefore assessable by inferential standards.
Intelligent Epistemology, §31
Verdict The epistemic is–ought is resolved constitutively rather than derived. The standards of inference bind outputs offered as inference, in the way the rules of proof bind a purported proof and preservation laws bind a purported homomorphism. An independent normative premise is needed only for whether to enter the inquiry, how hard to work at it, and how epistemic goods trade against others.
▸From the paper
An agent or process that presents an output as inference is bound, in that capacity, by the standards of Part I: explicit grounds, valid dependence, invariance under equivalent descriptions, and selection fixed by the problem.
In that capacity is the whole qualification. The theorem binds the output, not the agent's life. It says what a claim of inference has already committed to. No premise about what anyone ought to care about is needed.
▸The proof idea · Thm 31.1argued
Norms of inference
HypothesesAn agent or process that presents an output as inference.
The method theorem analyses what it is for an answer to count as following from its grounds.
An output satisfies that description exactly when every dependence it carries is carried by its stated grounds.
An output also presented as complete must additionally preserve every supported case and every supported distinction.
The standards are therefore internal to the claim being made, not added to it from outside.
A reasoned rejection of every inferential norm is itself offered as an inference and falls under the same standards.
Within the agent domain and consistency requirements stated in Intelligent Economics, bounded valued comparison against a reference forces the same exponential choice law, its score decomposition, its KL-regularised variational dual, and a canonical reversible relaxation.
Intelligent Systems — The Adaptive Closure · Inherited Result 4.2
Intelligent Economics reaches the same choice law from the other side, starting from bounded agents comparing options against inherited expectations rather than from what an answer may contain. Two derivations, different premises, one operator. The keystone paper imports both and asks what survives composition.
The paper does not claim the classical problems are finished. It
claims each has been pushed until it states its own remaining work,
and that the remaining work has five shapes. Each shape below carries
its answer type and the theorem that receives it.
(i)
Candidate generation
Inquiry enlarges the problem with a representation, hypothesis, or rule. Nothing inside a fixed candidate class can reach a candidate outside it.
Further computation reaches more of a full answer that was already fixed. Deductive closure is the normative object; the accessible subset is where a finite agent stands.
One fact at the base: consistent inference is possible. Everything above it is what that fact contains once each domain has named its objects. The principle adds nothing to any problem it meets. By adding nothing it lets each problem show exactly what it holds.
In plain words The whole book rests on something so small it is almost embarrassing to state: reasoning can be done consistently. That is it. Every law, every resolution, every boundary in these pages is what follows from taking that one fact seriously inside a properly stated question.
MU is the floor of this account. The floor is thin by design. Consistent inference is possible; nothing more is claimed at the base. Everything built above it is what that one fact, taken seriously inside each stated domain, turns out to contain. An answer offered as following from its grounds must draw on those grounds alone. Every answer-relevant dependence therefore lives in the problem. Equivalent descriptions agree. A complete report keeps the whole supported result.
Held to arithmetic, the discipline takes four forms. Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Under finite counting symmetry it is Shannon's maximum entropy. Revision is KL projection, carrying forward every feature the new constraints permit. Values and costs then govern choice, channels carry belief's connection to the world, and four world-facing conditions decide when that connection finds the truth. These are laws rather than customs. Within their stated domains, no consistent alternative exists.
What the structure says
What the problems became
Regress ends at the exhibited floor. Indifference returns a point, a family, or a certified boundary according to the structure present. Induction becomes a set of explicit projective and world-facing conditions with the price of each on the table. Goodman moves projectibility into representation. Gettier joins internal support to route integrity. Scepticism separates the self-defeating universal doubt from the tractable local kind.
Scientific realism reaches the quotient its experiments fix, and the limit theorems contribute their ceilings and empty classes as complete answers in their own right. A certified boundary is a finding. A lawful refusal is an answer with its reasons attached. Both are results.
What the structure says
What is left
Work with a shape, and only five of them. Generate a candidate, design a discriminating experiment, compute further, state the value at issue, or record the equivalence that survives. Each task is the same discipline carried into new ground. Nothing new is needed.
And the ground itself asks for nothing. It was there before the argument began, and every argument, including any brought against this one, stands on it.
Π ⟼ Answer(Π)the whole operator
The problem gives the subject matter. MU returns its complete answer. The structure present fixes the resolution: a point, an orbit, a wider family, or an empty requested class certified by proof.
Determinate Boolean plausibility is probability. Total informational change is relative entropy. Least-assuming completion is relative-entropy projection. Revision is KL projection. Within their stated domains, no consistent alternative exists.
One rule for what follows. Four quantitative laws. Every answer shape.
Intelligent Epistemology, the paper's closing line
Verdict Held to arithmetic, the discipline takes four forms and each is unique in its domain. Met by the classical problems, it resolves each at exactly the level its grounds support. What remains has five shapes. And the floor was never in question: every argument, including any brought against this one, stands on it.
▸From the paper
It was there before the argument began, and every argument, including any brought against this one, stands on it.
Conclusion · p. 77
The paper's last sentence before its closing display. Read it against Corollary 2.2 on page 4. There the denial supplies the witness. Eighty pages later the same move closes the book.
It was there before the argument began, and every argument, including any brought against this one, stands on it.
Afterword · The descendants
What this paper grounds.
Intelligent Epistemology settles what an answer may contain
when it must follow from stated grounds. Nothing below MU is assumed.
Intelligent Economics reaches the same operator from the other
end: bounded systems choosing under inherited expectations, priced by
the cost of departing from what they already carry, arriving at the
identical exponential form by a route that never mentions grounds.
Intelligent Systems imports two of the results below verbatim
and asks what survives composition, observation, scale, and time.
The First Principle carries the discipline in book voice.
Each object this paper settles, and the work that carries it
The First Principle ↗the same discipline in book voice, without the apparatus. Every section of this edition names the chapters that tell its part of it.
Not carriedthe one object this paper does not supply. Intelligent Systems makes generation a separate operator, priced by selection after the fact.
Nothing in the right-hand column is reproved downstream. Each work
names its import, cites the result, and carries it across. A claim that
cannot name where its structure entered is the claim this paper calls
smuggled. That is the whole test.