The Goalposts: Impossible-to-Ignore Standard
A precommitted burden of proof for deciding whether Constraint-Based Reconstruction is merely an interesting research method, a viable physical framework, or evidence strong enough that a skeptical expert must engage the result rather than dismiss it.
We have repeatedly crossed a finish line and then discovered that it was only a checkpoint. This document moves the finish line before the next play begins.
Contents
- The referee answers first
- Why this edition raises the bar again
- NASA-facing standard: open science is the floor, not the finish line
- Why our earlier goalposts were too close
- What “impossible to ignore” means—and what it does not mean
- The two-layer rule: comprehension is part of rigor
- The claim ladder: five different things we might mean by “it works”
- The reviewer panel: five people we must satisfy at once
- Before scoring: distinguish a fatal loss from a local loss
- The actual championship finish line
- The impossible-to-ignore finish line: twelve burdens, one frozen branch
- How a reviewer would decide whether the model is truly overconstrained
- Discovery-grade evidence: how high should the numerical bar be?
- What would count as a genuinely convincing prediction?
- The anti-self-deception protocol
- Championship burden I — Mathematical existence is not enough; the selected physical state must be viable
- Championship burden II — Forcedness must be proved before a precision mismatch can be interpreted
- Championship burden III — Parameter economy must be measured as effective flexibility, not marketing arithmetic
- Championship burden IV — The theory must be overconstrained after calibration
- Championship burden V — Cross-domain success must come from one mutually compatible branch
- Championship burden VI — Cosmology requires a native background and a selected state, not merely a successful transfer calculation
- Championship burden VII — The theory must compete with alternatives, not merely fit nature in isolation
- Championship burden VIII — Independent replay and a genuinely new prediction are the final proof-of-seriousness
- Championship burden IX — The theory must make a discriminating prediction, not just an accurate one
- Championship burden X — The result must survive reasonable formulation, scheme, and implementation changes
- Championship burden XI — A surprising cross-domain link must survive as one calculation
- Championship burden XII — The decisive argument must be explainable without weakening it
- The hostile reviewer’s ten objections—and what would answer each one
- How much success is enough? Avoiding a new arbitrary goalpost
- The skeptical-reviewer forced-choice theorem
- Sector finish line — Geometry and gauge structure
- Sector finish line — Flavor, masses, and mixing
- Sector finish line — Quantum consistency, state space, and boundary admissibility
- Sector finish line — Granularity, records, and the first-separability clock
- Sector finish line — Primordial correlations and the inflation-replacement claim
- Sector finish line — Full cosmological history
- Sector finish line — What would make “unification” more than a common language?
- How I would update belief as the tournament progresses
- The public-language goalposts
- How repairs are allowed without turning falsification into an impossible game
- Formal same-ruler audit — proving that theory and experiment are actually the same object
- Uncertainty, truncation, and scheme dependence — a prediction is a distribution, not a favorite number
- Candidate-search accounting — how to pay for ideas tried before the winner appeared
- The prospective prediction protocol — exactly how we should run the next championship test
- The joint cross-domain score — one scoreboard for the whole theory
- A simulated final referee report: what I would write if the manuscript arrived today
- The hard-stop constitution — results that end the current branch rather than create another appendix
- The chronology standard — proving that a prediction existed before the answer
- The master reviewer matrix — claim, proof, falsifier, repair, and promotion
- Tournament rules v2 — a story humans can follow without lowering the physics bar
- The championship game plan — the exact sequence that maximizes information and minimizes self-deception
- Evidence replay: put the current branch against the new goalposts
- The clock audit
- The free-parameter combine
- The microscopic clock
- The local-correlation quarterfinal
- The hidden-global-mode semifinal
- The state-selection comeback
- The no-fudge scalar-spectrum test
- The vacuum-stability red-zone match
- The flavor prediction upset
- The baryogenesis endurance round
- The dark-sector possession test
- The home-field cosmology test
- The referee audit
- The championship conditions
- Would I call the theory correct today?
- The reviewer’s current scorecard
- The order of operations: where the goalposts say we should spend effort
- What would actually convince me?
- What would not convince the reviewer—even if it looks impressive
- The NASA-facing championship evidence packet
- The prediction packet we should freeze before the next major test
- Independent reviewer protocol
- The no-moving-goalposts decision tree
- The only claim worth fighting for
- Final whistle: where the goalposts really belong
- Appendix A — Frozen burden-of-proof matrix
- Appendix B — Parameter and flexibility constitution
- Appendix C — Source integrity ledger
- Appendix D — Reviewer checklist before any future “we are done” claim
The referee answers first
If I were reviewing this program for a serious technical audience, I would not ask whether the project has solved every philosophical question, derived every constant from nothing, or eliminated every conventional input. Those are unrealistic standards. I would ask something both harder and more useful: has one frozen mathematical branch survived enough independent, prospectively defined tests that the remaining probability of coincidence, flexible fitting, bookkeeping artifacts, or hidden inconsistency is small?
The answer requires a different finish line from the gate board. A theory can be impeccably governed and still be false. A theory can contain measured anchors and still be scientifically powerful. A theory can fail one auxiliary ansatz without losing its core. The task is to distinguish these cases before results arrive.
We may call the current branch strongly supported within its tested domain only after it clears all branch-fatal consistency/stability tests, carries a complete and honest parameter ledger, produces an overconstrained vector of held-out predictions from one mutually consistent branch, survives comparison with relevant benchmark theories under a common statistical ruler, passes destructive controls, and is independently reproduced from frozen artifacts. Until then, stronger words such as “correct,” “unified theory,” or “reconstruction of nature” remain hypotheses rather than earned conclusions.
This is intentionally more demanding than “33 gates resolved,” “the equations are internally coherent,” or “several outputs are numerically close.” It is also more attainable than the impossible demand to prove a physical theory true for all regimes forever.
Why this edition raises the bar again
The purpose of this document is not to help us win an argument. It is to make it difficult for us to fool ourselves.
If we want a skeptical scientist to change their mind, it is not enough to show that our equations can reproduce known facts. A flexible framework can often do that. We need to put the theory in situations where it could clearly lose, freeze the rules before seeing the answer, and then let nature decide.
The finish line is therefore not “we closed the gates.” The finish line is “a hostile expert can no longer dismiss the result without pointing to a specific error, hidden degree of freedom, data problem, or failed prediction.”
The evidentiary target is an adversarially reproducible, overconstrained, prospectively scored theory comparison. Let H_CBR denote a frozen theory version, B_j reasonable benchmark models, D_cal the declared calibration set, and D_hold held-out data. Promotion is permitted only when the predictive distribution p(D_hold|H_CBR,D_cal) is generated without target-dependent repair, the effective flexibility of H_CBR is charged, and the result survives independent replay and benchmark comparison.
No theorem can make an empirical theory literally irrefutable. The correct high bar is that refutation must engage a specific falsifiable object: an equation, assumption, provenance link, implementation, same-ruler map, or prospective datum.
NASA-facing standard: open science is the floor, not the finish line
If the intended audience includes NASA scientists, the evidence package should meet the reproducibility and transparency norms they already expect before we ask them to consider the physics.
NASA scientists should not have to trust us. They should be able to download the inputs, run the calculation, inspect the assumptions, and see the same result. If a crucial number exists only in our prose, the burden has not been met.
NASA Science currently describes open science in terms of transparent, available, reproducible, and collaborative scientific processes. Its Science Information Policy requires open availability of publications and emphasizes sharing data and software; NASA guidance also explicitly connects software and archived input/configuration files to reproducibility. NASA scientific-integrity guidance emphasizes peer review, disclosure of assumptions and biases, open sharing of methods/results where possible, and honest reporting.
These are process requirements, not evidence that NASA endorses any theory. We therefore use them as a minimum reproducibility floor and impose a substantially stronger theory-validation burden on top.
Why our earlier goalposts were too close
Our earlier process repeatedly made the same category mistake: it treated completion of a local burden as completion of the global claim. Closing a gauge-group gate does not prove flavor. Reconstructing flavor does not prove vacuum stability. Deriving a first-record time does not prove primordial perturbations. Selecting a lawful state does not prove that the state has the right spectrum. A governance terminal says the scoped question has reached an accepted endpoint; it does not automatically say that nature selected the whole branch.
There are four recurrent ways a finish line moved after we crossed it:
- Scope promotion: a theorem about one sector was unconsciously promoted to the entire theory.
- Terminal promotion: a CLOSED-NEGATIVE, MEASURED-ANCHOR, or CERTIFIED-IRREDUCIBLE terminal was mentally counted like a derived empirical success.
- Calibration promotion: a number inherited from an earlier fit was later experienced as though it had been predicted from geometry alone.
- Mechanism promotion: identifying a mathematically possible mechanism was treated as equivalent to deriving the unique mechanism chosen by the frozen theory.
The cure is not pessimism. The cure is a hierarchy of claims with explicit promotion rules. Every higher rung requires all lower-rung obligations plus genuinely new evidence.
What “impossible to ignore” means—and what it does not mean
We cannot make a scientific idea impossible to refute. If we did, it would no longer be scientific. What we can do is make the evidence so clean that disagreement has to become specific.
A strong result is one where a reviewer cannot reasonably say “you probably tuned it” or “there must be another branch” because the tuning budget, branch grammar, chronology, code, and held-out targets were frozen before the test.
Define a decisive claim packet P=(H,\Theta,G,C,S,U,F), where H is the hashed theory version, \Theta the complete effective parameter/function ledger, G the candidate/branch grammar, C the calibration contract, S the scoring rule, U the uncertainty/scheme specification, and F the falsifier/stop-rule set. The packet is timestamped before target access.
A post-test objection is technically relevant only if it attacks one of these frozen objects, the data, the implementation/replay, or the logical implication from them. This does not make the theory true by decree; it prevents vague post-hoc dismissal from substituting for scientific criticism.
The asymmetry we want
It should remain easy for the theory to lose and hard for it to win. One core-forced contradiction can reject a branch. By contrast, many successful reconstructions are required before a strong correctness claim is earned.
The two-layer rule: comprehension is part of rigor
A proof burden that only its authors can understand is weaker than it looks. Every decisive result in this program should therefore be presented twice.
The first layer answers: What are we claiming, why should I care, what would make it wrong, and what happened? A scientifically literate reader should be able to follow the logic without reconstructing every tensor contraction.
The simple layer is not allowed to change the claim. It can remove notation, but not assumptions, caveats, or failure conditions.
The technical layer must contain the exact mathematical object, domain, assumptions, derivation or executable recipe, provenance, uncertainty propagation, controls, same-ruler observable map, and terminal status. If the simple and technical layers imply different claims, the technical layer controls and the communication gate fails.
The six-line claim format
| Line | Required content | Why it exists |
|---|---|---|
| 1 | Plain-language claim | Prevents rhetoric from hiding the proposition. |
| 2 | Exact mathematical claim | Defines what was actually proved or predicted. |
| 3 | Inputs and assumptions | Shows what the result inherited rather than derived. |
| 4 | Calculation / witness | Makes the result reproducible. |
| 5 | How it can fail | Keeps the claim falsifiable. |
| 6 | Status | PASS, FAIL, OPEN, CALIBRATED, or CONDITIONAL—never blended. |
The claim ladder: five different things we might mean by “it works”
Level 1 — Useful research method
The constraint-first method produces reproducible calculations, finds contradictions, improves building blocks, and can reject favored hypotheses. It need not yet identify the correct fundamental theory.
Promotion condition: the method must demonstrably catch known defects and produce negative results when warranted.
Level 2 — Internally viable physical branch
The selected Shape/Rulebook/Actor branch is mathematically well-defined over its claimed regime, has a lawful physical vacuum/background, no uncancelled anomalies or ghosts in scope, and no confirmed forced contradiction with established observations.
Promotion condition: all branch-fatal consistency and empirical red-zone tests clear.
Level 3 — Predictive theory
After paying every fitted parameter and branch choice, the theory predicts additional observables that were not used to build or choose it. The held-out data overconstrain the effective parameter dimension.
Promotion condition: prospective freeze, honest complexity, same-ruler comparison, and successful held-out vector.
Level 4 — Strongly supported unification candidate
One frozen branch explains multiple traditionally separate domains—particle structure, precision phenomenology, quantum/state structure, and cosmology—with a common parameter/provenance ledger, while remaining competitive with domain benchmarks.
Promotion condition: cross-domain joint success without incompatible branch changes or domain-specific arbitrary functions.
Level 5 — Strong evidence the framework captures real underlying structure
The theory makes at least one genuinely distinguishing prospective prediction, survives independent implementation/review, and continues to predict new data better than flexible alternatives after the original development team stops adjusting the core.
Promotion condition: independent replication plus successful novel prediction. Even here, “correct” means strongly supported within tested scope—not metaphysical certainty.
The reviewer panel: five people we must satisfy at once
Reviewer A — Mathematical physicist
“Show me the actual dynamical object. Is the action/function space well-defined? Are the constraints first-class where claimed? Are gauge directions separated from physical negative modes? Is the selected vacuum stable or controlled-metastable? Is the initial-boundary problem well posed? Where do you rely on an axiom or external theorem rather than a derivation?”
Reviewer B — Particle phenomenologist
“Count every fitted real. Show me the scale, scheme, running, matching, and uncertainty propagation. If a quantity is called forced, compare it to precision data without re-fitting. A 14σ discrepancy is not a footnote. Tell me whether it kills a one-angle ansatz, a sector, or the entire branch.”
Reviewer C — Cosmologist
“Do not give me a primordial spectrum by declaration. Derive the background, the physical perturbation variables, the state-selection rule, and the transfer history. Then carry the same solution through BBN, recombination, CMB, BAO, growth, and age. Tell me which cosmological numbers were fitted and remove them from the prediction score.”
Reviewer D — Statistician / model-selection skeptic
“How many effective degrees of freedom does this construction have, including branch choices and free functions? When were they chosen? What data were visible? Show me out-of-sample predictive scores against reasonable benchmarks, not only a list of individually attractive numbers.”
Reviewer E — Independent reproducer
“Give me the frozen inputs, hashes, code, branch manifest, and falsifiers. I will implement the decisive calculation independently. If I cannot reproduce it without asking the builders which normalization they intended, the result is not yet strong evidence.”
The theory does not get to choose its friendliest reviewer. A claim of deep unification must survive the intersection of all five burdens.
Before scoring: distinguish a fatal loss from a local loss
A central source of moving goalposts is failing to decide what a negative result actually kills. The rule should be structural, not emotional.
| Negative result | What it kills | When it kills the whole frozen branch |
|---|---|---|
| Auxiliary ansatz fails | The ansatz | Only if the ansatz was proven uniquely forced by immutable core constraints. |
| One Actor realization fails | That Actor realization | If the admitted Actor grammar was exhaustively closed and no alternative lawful realization exists. |
| Forced observable contradicts data | The forcing chain | If the same-ruler map and experimental record are sound and no upstream ambiguity remains. |
| Physical vacuum unstable | The selected background | If the instability is in a physical direction and no already-admitted prospective completion stabilizes it. |
| Cosmology kernel not derived | Cosmological completeness claim | Not necessarily particle branch; whole unification claim fails until a lawful state/background derivation exists. |
| Free parameter needed | Zero-parameter claim | Not the theory, if the parameter is openly paid and held-out predictions remain. |
The actual championship finish line
A reviewer would realistically require the following eight burdens to be cleared simultaneously before treating the current branch as strongly supported.
- Mathematical viability: a well-defined effective/parent dynamical problem in every claimed regime, with anomaly/gauge consistency, lawful physical state space, and no uncontrolled ghost/tachyon/negative-norm sector.
- Stable selected background: the actual frozen vacuum/cosmological saddle is stable or quantitatively metastable after the corrections the theory itself admits.
- Complete provenance and complexity ledger: every measured anchor, fit, injected normalization, discrete branch, free function, and external theorem is typed and charged before prediction scoring.
- No known forced empirical contradiction: precision observables claimed as predictions agree within a preregistered uncertainty model; any surviving high-significance contradiction is adjudicated as local or branch-fatal by the rule above.
- Prospective overconstraint: after calibration, the held-out observable vector contains more independent information than the remaining effective flexibility and is evaluated without re-fitting.
- Cross-domain coherence: one branch and one parameter ledger survive particle physics, quantum/state tests, and cosmological history. Domain-specific incompatible branches do not count as unification.
- Distinctive prediction: at least one outcome that reasonable competitor theories would not generically have predicted is frozen before the relevant datum is opened or measured.
- Independent replay: a technically competent outsider reproduces the decisive chain from frozen artifacts and obtains the same verdict.
Nothing in this list requires deriving ℏ, MPl, or every Standard Model input from nothing. Nothing requires beating every established theory on every legacy datum. It requires something more scientifically relevant: the theory must become harder to flex than the evidence it successfully predicts.
The impossible-to-ignore finish line: twelve burdens, one frozen branch
The championship goalpost is deliberately higher than the earlier document. Clearing all gates is not sufficient; clearing all twelve burdens below is the minimum for a strong “this framework is probably capturing real physics” claim.
| # | Burden | Non-negotiable evidence |
|---|---|---|
| 1 | Physical viability | Stable/metastable lawful background; no unresolved fatal quantum inconsistency. |
| 2 | Forcedness | Proof that decisive outputs are consequences of the frozen branch rather than optional ansätze. |
| 3 | Full flexibility accounting | Parameters, functions, branches, priors, search attempts, and repairs charged. |
| 4 | Overconstraint after calibration | Substantial held-out information remains after every paid calibration. |
| 5 | One branch across domains | No incompatible sector-specific winners. |
| 6 | Native cosmology | Background + state + matter ledger derived or explicitly calibrated before cosmology scoring. |
| 7 | Benchmark competition | Same data, same nuisance treatment, complexity-aware comparison. |
| 8 | Independent replay | Independent implementation reproduces decisive results. |
| 9 | Discriminating prospective prediction | A frozen prediction on which reasonable alternatives differ materially. |
| 10 | Robustness/invariance | Headline result survives nonphysical formulation/scheme/implementation choices. |
| 11 | Cross-domain inheritance | At least one upstream object genuinely constrains independent domains. |
| 12 | Human-auditable explanation | Simple and technical layers agree; assumptions/calibration/falsifiers visible. |
How a reviewer would decide whether the model is truly overconstrained
Counting “outputs” is not enough because outputs can be correlated and free functions can carry enormous effective flexibility. The correct object is the predictive distribution for a held-out data vector yH after calibration data yC have already been consumed.
For a deterministic branch with a frozen best calibration and approximately Gaussian measurement covariance ΣH, a simple diagnostic is
but the reviewer should also compare predictive log score or marginal likelihood against relevant benchmark models. The goalpost is not an arbitrary χ² threshold chosen after inspection. It is: acceptable absolute predictive fit plus competitiveness against benchmarks after paying complexity.
For local identifiability, let J be the Jacobian of independent observables with respect to the paid parameter vector. If the calibration and held-out data constrain only the same low-rank combination of parameters, the apparent number of successes exaggerates evidence. The Fisher information
must have the rank expected from the claimed parameter count, and held-out observables should probe directions not already exhausted by calibration. This is why “five parameters and fifty plotted points” is not automatically impressive; fifty strongly correlated points may carry only a handful of independent modes.
A calibrated datum pays for its parameter and then leaves the scoreboard. Evidence begins with the predictive consequences that remain after that payment.
Discovery-grade evidence: how high should the numerical bar be?
There is no universal magic sigma or Bayes-factor threshold for all of physics. The defensible rule is to choose the scoring system before seeing the target and make the headline threshold intentionally difficult.
For an ordinary compatibility claim, being within error bars may be enough. For a claim that an established picture is fundamentally incomplete, compatibility is nowhere near enough. We need evidence that is hard to get by chance, hard to get by tuning, and hard for reasonable alternatives to reproduce.
For each decisive test, preregister a domain-appropriate evidentiary criterion. Examples of deliberately high conventions include a discovery-grade frequentist threshold for a clean new effect, or a log Bayes factor above 5 (roughly >150:1) for a model comparison under declared priors, accompanied by sensitivity analysis and independent replication. These numbers are not laws of nature; their scientific value comes from being specified before the result and accompanied by multiplicity/search corrections.
For the overall theory claim, no single p-value or Bayes factor is sufficient. The relevant object is the joint predictive evidence across independent held-out domains, with calibration and shared systematics accounted for.
For a public claim that the framework overturns or materially supersedes an established physical account, require: no hard-stop failures; decisive prospective evidence under a frozen scoring rule; independent replay; and a joint complexity-aware comparison that remains favorable across reasonable prior/scheme choices.
What would count as a genuinely convincing prediction?
Legacy reconstruction is useful but intrinsically vulnerable to selection effects: the model builder knows what nature already looks like. The strongest evidence comes from a prediction whose target was not available to guide the branch.
A reviewer would rank predictions roughly as follows:
| Evidence class | Example form | Weight |
|---|---|---|
| Re-description | Known datum encoded directly as a measured anchor or category pin | Necessary for calibration, not predictive evidence |
| Retrodiction after broad model search | A known mass/angle reproduced after candidate exploration | Weak-to-moderate; selection history matters |
| Frozen held-out retrodiction | Observable withheld while a neighboring sector is built | Moderate-to-strong if the firewall is credible |
| Cross-domain prediction | Particle-derived scale fixes an early-universe quantity without cosmological fitting | Strong if provenance is independent and downstream result is nontrivial |
| Prospective novel prediction | Numerical/qualitative outcome timestamped before new measurement | Highest weight |
The final theory does not need dozens of sensational novel predictions. One or two genuinely distinguishing prospective successes, combined with a broad held-out legacy score and no fatal contradictions, would change the evidentiary picture far more than another twenty post-hoc matches.
The anti-self-deception protocol
The project already contains many of the right controls; the final reviewer standard makes them mandatory at the theory level.
- Freeze hashes before target access. Record the exact Shape/Rulebook/Actor manifest, code, parameter ledger, candidate grammar, and prediction packet.
- Separate builders from reviewers. The reviewer may identify defects but may not repair them in the same pass. This prevents the audit from unconsciously searching only for fixable errors.
- Maintain a failure registry. Every killed branch remains visible. Deleted failures create survivorship bias.
- Use negative controls. Wrong geometry, permuted labels, wrong state, wrong sign, wrong parity, wrong dimension, or deliberately perturbed parameters should fail in predictable ways.
- Run ablations. Remove Shape, Granularity, State Selection, or a proposed global mechanism and measure what predictive structure disappears.
- Charge branch search. If ten candidate mechanisms were tried and one worked, the selection cost belongs in the evidence assessment even if no continuous parameter was fitted.
- Propagate uncertainties. A prediction with uncertain upstream anchors is a distribution, not a point. Same-ruler significance uses the full propagated covariance.
- Independent code path. At least one decisive calculation should be reimplemented without sharing the original implementation details beyond the mathematical specification.
- Prediction lockbox. For prospective tests, publish a cryptographic hash or timestamped artifact before the measurement is known.
- Predefine stop rules. A branch-fatal contradiction causes rejection even if the rest of the paper is beautiful.
This protocol is intentionally asymmetric: it is easier to keep exploring a theory than to earn the right to call it correct.
Championship burden I — Mathematical existence is not enough; the selected physical state must be viable
Before asking whether the theory predicts nature, we first ask whether the physical state it uses actually exists and is stable enough to support the calculation. If the assumed vacuum falls apart, downstream successes do not rescue it.
Let \bar\Phi be the selected background and \Gamma[\Phi] the relevant effective action. Eligibility requires the gauge-reduced physical Hessian \delta^2\Gamma/\delta\Phi_{phys}^2|_{\bar\Phi} to have the correct signature after all required loop/Casimir/boundary contributions and scheme uncertainties are included. A negative physical eigenvalue is not dismissed unless a controlled metastability calculation demonstrates an acceptable lifetime.
The first reviewer burden is often underestimated because a beautiful action can make a model feel more complete than it is. A theory does not earn physical viability merely by writing a compact manifold, an effective action, or a stationary point. The actual state used by every downstream prediction must exist in the physical configuration space and be dynamically admissible.
The reviewer therefore asks for a chain that begins before any phenomenology:
The word physical matters. Gauge zero modes, coordinate artifacts, Lagrange-multiplier directions, and redundancies must be removed before interpreting signs of the Hessian. Conversely, a negative eigenvalue cannot be dismissed merely because it appears in a large unreduced matrix. The document must publish the reduction map that identifies the physical tangent space and then evaluate the second variation there.
At tree level, the minimum burden is straightforward. If Φ denotes the background fields and Φ̄ the selected solution, then the first variation must vanish on all allowed physical variations,
and the quadratic form
must have no uncontrolled negative physical direction for a claimed stable vacuum. If the theory only requires metastability, the negative or tunneling sector must instead yield a quantitatively acceptable lifetime from a specified decay/bounce calculation. “Perhaps loops fix it” is not a stability result.
The one-loop burden is also easy to state even when hard to execute. Every correction used to stabilize a tree-level saddle must come from the already admitted field content and regularization/renormalization grammar. The sign of the relevant curvature of the effective potential must be scheme-controlled to the precision needed for the claim. If the sign can be flipped by an un-fixed orientation bit, subtraction prescription, or renormalization condition, then the state is not yet predictively stable from frozen data.
A reviewer would also demand a domain statement. It may be enough for a compactification vacuum to be metastable on timescales enormously longer than the cosmological history of interest; absolute global stability is not always required. But the lifetime calculation must use the same physical potential and fields that define the rest of the theory. A nominally stable reduced toy potential cannot certify an instability that reappears when the omitted shape mode is restored.
A complete physical fluctuation analysis around the selected branch shows either (a) nonnegative physical spectrum with only understood symmetry zero modes, or (b) a controlled metastable decay rate safely outside the claimed physical history. Loop corrections and scheme dependence are bounded tightly enough that the sign/lifetime cannot be chosen by convention.
A reproducible negative physical mode remains and every proposed stabilization requires a new uncharged term, an after-the-fact scheme choice, or a new branch selected because it repairs the instability. If the unstable background is constitutive of the frozen branch, that is branch-fatal.
This is why the present vacuum-stability item is not a bureaucratic residual. It sits at the very first championship burden. If the field configuration assumed by the particle and cosmology calculations is not a viable physical state, successes computed on that configuration do not rescue it.
Championship burden II — Forcedness must be proved before a precision mismatch can be interpreted
A bad prediction is only fatal if the theory truly had no freedom to choose a different answer. We must prove what was forced before using either success or failure as evidence.
For observable O, define the admissible candidate set \mathcal A(H). “Forced” means the pushforward O_*\mathcal A(H) is a singleton (or a preregistered narrow distribution from declared uncertainties). If multiple lawful branches produce materially different O, then the observable tests the selection rule/auxiliary ansatz rather than the immutable core.
The second burden is to distinguish a falsified theory from a falsified ansatz. This is crucial because an honest program can otherwise oscillate between overclaiming and overreacting: first it calls a convenient relation “forced,” then when data disagree it calls the relation merely illustrative.
Before opening the target observable, the candidate grammar must be stated. Let C be the set of lawful constructions under the frozen Shape, Actor, Rulebook, and building-block constraints. A result O=O★ is genuinely forced only if every surviving candidate in the completed grammar yields the same output within the declared uncertainty:
If only one convenient ansatz c₀ has been evaluated, the correct label is “prediction of c₀,” not “prediction of the theory.” Exhaustion can be analytic—by a classification theorem—or computational, provided the search space, priors, discretization, and stopping rules are explicit. Merely failing to think of an alternative is not an exhaustion proof.
Once forcedness is established, same-ruler comparison becomes uncompromising. If the theory predicts a renormalized quantity at scale μ in a particular scheme, it must be evolved and projected to the same object reported experimentally. If mixing conventions, absolute values, marginalization choices, or correlations differ, the mismatch cannot be scored until the map is fixed. But once the ruler is genuinely the same, a large discrepancy is evidence.
The review should use the full covariance rather than isolated “sigma” arithmetic whenever observables are correlated. For a predicted vector μ and measured vector y,
is a cleaner measure than summing marginal pulls. The important principle is not the conventional 5σ number by itself. It is that a high-significance disagreement must not be downgraded after the fact unless the theory-level forcing statement was wrong. In that case the correct correction is to retract the forcedness claim, explain why the grammar was incomplete, and pay for the expanded grammar in the complexity ledger.
A prospective repair is scientifically legitimate. Suppose the one-angle flavor ansatz fails and an independent structural analysis—performed without reading the discrepant observable—shows that the frozen geometry necessarily supplies a second invariant phase. Adding that phase is not cheating merely because it repairs the data; it becomes a new theory version whose complexity is paid, whose target is removed from the prediction score, and whose other predictions are newly tested. What is forbidden is searching phase structures while watching |Vtd| and then presenting the selected structure as though geometry uniquely demanded it.
For every high-value “forced prediction,” the candidate grammar is demonstrably exhausted or the claim is narrowed to the evaluated ansatz. Same-ruler uncertainty propagation is reproducible. No confirmed core-forced precision discrepancy remains.
A core-unique prediction remains strongly inconsistent with established data after independent replay. If repair requires changing the frozen core or selecting new degrees of freedom after opening the target, the current branch loses.
This burden is what makes the ~14σ flavor item so decision-relevant: not because 14 is a magical number, but because a discrepancy that large forces us to settle the logical status of the one-angle mechanism. The correct goalpost is not “explain the tension”; it is “prove whether the tension belongs to an optional ansatz or to the frozen branch, then score it accordingly.”
Championship burden III — Parameter economy must be measured as effective flexibility, not marketing arithmetic
We count every way the theory could have been bent toward the answer—not just the parameters written in one equation.
Effective flexibility includes continuous parameters, free functions, basis/branch choices, priors, stopping rules, model-selection tries, data-dependent transformations, threshold choices, and repair opportunities. Complexity must be scored at the level of the entire search procedure, not the final surviving formula.
A unification theory is allowed to have parameters. The problem is not fitting; the problem is claiming explanatory compression when the real flexibility has been hidden in normalizations, branch searches, boundary functions, or convention choices.
The reviewer therefore asks for a complete generative ledger. Every quantity capable of changing the likelihood of an observable is typed before evaluation. Continuous parameters are obvious, but the less obvious flexibility can dominate: discrete geometry choices, candidate topologies, sign selections, operator bases, truncation orders, threshold normalizations, priors, and state-kernel functional forms.
A free function is especially dangerous. Writing w(z), K(k), or F(Φ) as one named object does not make it “one parameter.” If the function is represented by m independent basis coefficients over the range probed by data, its effective dimension is of order m, reduced only by genuine prior/regularization information that was fixed independently of the target. A Gaussian-process prior likewise carries an effective complexity determined by its kernel and hyperparameters; it is not free explanatory power.
Discrete search also costs evidence. If N roughly comparable branches were examined using the target and the best was reported, the evidentiary effect resembles a look-elsewhere factor. Exact Bayesian accounting would integrate over the prior branch weights,
rather than score only the winning b. A simpler project ledger can at least publish the number and nature of serious alternatives tried and prevent the winner from being described as uniquely forced unless the alternatives were independently eliminated.
The goalpost is therefore not “fewer parameters than the Standard Model” in a naive count. It is competitive predictive compression: after paying the effective flexibility that was actually used, does one parameter/structure package account for more independent data than reasonable benchmarks?
Information criteria can be useful diagnostics when regularity assumptions hold:
but the final paper should not fetishize either. For hierarchical or singular models, out-of-sample predictive log score or full marginal likelihood is safer. The principle is invariant: flexibility is paid before evidence is claimed.
A versioned ledger lists measured anchors, fitted reals, discrete branch choices, free functions with effective dimension, nuisance parameters, and external floors. The model remains overconstrained and benchmark-competitive after this honest accounting.
The apparent predictive economy disappears when hidden normalizations, searched branches, or functional freedom are charged. This would not make the model mathematically false, but it would defeat the strong unification/compression claim.
This burden also prevents a common psychological error: “we only added one building block.” A building block that introduces an arbitrary function can be far more flexible than ten scalar parameters. The score follows the degrees of freedom, not the vocabulary.
Championship burden IV — The theory must be overconstrained after calibration
After we use some observations to set allowed knobs, there must still be many independent facts left for the theory to get right. Otherwise we have calibrated a model, not tested it.
Partition data into D_cal and D_hold. The predictive score is evaluated only on p(D_hold|D_cal,H). A theory is meaningfully overconstrained when the effective information in held-out observables substantially exceeds the effective fitted flexibility and the held-out residual structure shows no systematic domain-dependent repair.
A calibrated theory can fit perfectly and still explain almost nothing. The championship standard begins only after calibration is over.
Let θ contain the paid cosmology- or particle-facing parameters, and let DC be the calibration data. The theory is frozen after constructing p(θ|DC). The held-out evidence DH is then evaluated through the posterior predictive distribution, not by re-optimizing θ:
The simple rule “one datum per parameter” is a useful minimum firewall but not a complete statistical definition because data and parameters can be correlated. A better test is whether the held-out observables probe independent combinations of the theory. The singular values of the whitened Jacobian Σ−1/2J show how many directions in observable space the parameters can readily move. If held-out residuals lie almost entirely in those same directions, the test is weak even if the plot contains hundreds of points.
Conversely, a small number of observables can be highly decisive if they test orthogonal structural consequences. Examples include a forbidden particle, a discrete charge ratio, a sign, a selection rule, a zero, or a spectral feature at a location fixed by an upstream scale. These can carry more evidentiary weight than a broad smooth curve whose amplitude and tilt were fitted.
The reviewer should also demand calibration swaps. If parameter A is calibrated to observable X and predicts Y, repeat the analysis calibrating to Y and predict X. The theory should remain coherent within propagated uncertainty. A model that works only for one privileged calibration convention may be hiding an identifiability problem.
Leave-one-domain-out tests are even stronger. Build the particle sector without cosmology, then ask cosmology. Or freeze the geometry/flavor construction without one precision observable and predict it. This is the closest legacy-data analogue of a truly prospective experiment.
After every paid calibration datum is removed, the remaining data vector contains independent structural information and is predicted without re-fitting. Calibration swaps and leave-one-domain-out replays preserve the branch.
Good agreement depends on continually reusing observables to select branches or update parameters, or the held-out data probe no directions beyond those already used in calibration. In that case the model may be a good fit but not yet a strong prediction engine.
This is the mathematical version of the sports metaphor: we do not count the warm-up points used to set the shooting sights. The scoreboard starts after the sights are locked.
Championship burden V — Cross-domain success must come from one mutually compatible branch
The same version of the theory has to win all games. We cannot use one branch for particle physics and a different incompatible branch for cosmology and then call the combination a unified theory.
There must exist a single parameter/branch assignment heta^*,b^* in the globally admissible model such that all sector likelihoods and hard constraints are simultaneously satisfied. Sector-wise optima ( heta_i,b_i) do not establish unification unless they are restrictions of the same global solution.
The word “unification” creates a special burden. It is not enough for one geometry variant to fit gauge structure, another to fit flavor, and a third to support cosmology. The same frozen branch must survive the interfaces between domains.
The reviewer therefore asks for a dependency graph rather than a list of papers. Nodes are derived objects—radii, spectra, couplings, vacuum expectation values, state kernels, background histories. Directed edges identify which object feeds which calculation. Every node has one authoritative value/distribution and one provenance class. If two papers silently use incompatible values of the same node, the theory has not unified those papers.
Cross-domain propagation is where the strongest evidence can appear because it denies the model opportunities to re-fit. A particle-physics scale that was fixed before cosmology and then determines an early-universe time is interesting precisely because cosmology did not set the scale. But its evidentiary label must preserve the particle-side calibration provenance. “Cross-domain consequence of an independently calibrated scale” is strong and honest; “parameter-free prediction” would be too strong if the scale was fitted upstream.
The same rule applies in the opposite direction. A cosmological state-selection principle that later constrains particle initial conditions could be powerful, but only if it was not introduced after a particle anomaly appeared. Cross-domain arrows must respect chronology.
Interface consistency also includes units, renormalization scales, frames, and effective descriptions. A 13D geometric parameter may compile to a 4D effective coupling only through a specified matching map. An object that is dimensionless in one normalization cannot be numerically compared to a different convention because both are called “mixing angles.” The assumption ledger’s same-ruler rule becomes an interface theorem here.
A single versioned dependency graph carries one branch through particle physics, state selection, and cosmology. Shared quantities have one authoritative provenance and uncertainty. Cross-domain predictions use frozen upstream objects without target-domain re-fitting.
Different domains require incompatible branch choices, contradictory values of shared quantities, or target-specific functions/normalizations. The individual domain models may remain useful, but the unification claim fails.
This burden is why the final evidence bundle should be judged jointly. Five successful papers that cannot coexist in one branch are not five pieces of evidence for one theory.
Championship burden VI — Cosmology requires a native background and a selected state, not merely a successful transfer calculation
For cosmology, we must derive the universe the theory evolves—not import the standard expansion history and only decorate it with our mechanism.
The native background is obtained by varying the compactified/effective action on the cosmological ansatz, while the primordial state is selected by an upstream boundary/admissibility principle. Only then may the resulting background and state be propagated through BBN, recombination, CMB, BAO, growth, and age. Importing a best-fit H(a) or target-shaped primordial kernel is calibration, not native cosmology.
Cosmology is the easiest place to accidentally move the goalposts because the downstream standard machinery is so powerful. If we import H(a), Ω values, and a primordial spectrum, then a Boltzmann solver can reproduce a great deal. That validates the solver, not the proposed origin theory.
The native cosmology burden has four logically separate layers:
- Background: derive the homogeneous solution from the compactified parent/effective action and matter ledger.
- Physical perturbations: reduce scalar/vector/tensor fluctuations to the gauge-invariant dynamical variables and their quadratic operators.
- State selection: use a prospective boundary/admissibility principle to determine the quantum state or finite admissible class.
- History transfer: propagate that background and state through thermal history to observables.
The background begins from a variational statement, schematically
with the Hamiltonian constraint and conservation laws satisfied. The theory must publish which stress-energy terms are genuinely derived and which are calibrated. If dark matter or vacuum energy is inserted as an arbitrary function merely to reproduce H(a), then the framework has not predicted the background.
The perturbation state is a distinct burden. A local boundary condition at finite time generally produces a kernel analytic in k² near k=0; the previous state-selection analysis showed why that cannot simply be declared to equal the required nonlocal primordial covariance. A local parent theory can nevertheless induce a nonlocal boundary kernel through its solved bulk history. The correct object is the on-shell/Dirichlet-to-Neumann Hessian. That route is valuable precisely because the nonlocality is derived from bulk propagation rather than inserted to match ns.
Only after background and state are frozen should the theory compute primordial amplitude/tilt/running, tensor power, non-Gaussianity, isocurvature, BBN abundances, CMB spectra, BAO distances, matter power/growth, and age. Any observable consumed to calibrate a remaining scalar parameter is removed from the prediction score.
The parent/effective action produces a viable cosmological saddle; SSBA selects the state without cosmological target loading; the induced spectrum and matter ledger are then propagated through a single thermal/perturbation history that survives multiple held-out cosmological probes.
A free primordial kernel, arbitrary H(a), or dark-sector function must be chosen from the cosmology data. That may yield a phenomenological cosmology, but it does not close the proposed native reconstruction.
This goalpost finally stops the repeated pattern “we derived the next boundary, therefore cosmology is almost done.” Cosmology is done only when the complete generative chain exists.
Championship burden VII — The theory must compete with alternatives, not merely fit nature in isolation
Matching nature is not enough. We must ask whether a simpler or established model explains the same evidence with fewer assumptions or better predictions.
Predeclare benchmark models and score them with the same observables, nuisance treatment, and data partition. Use a complexity-aware criterion—preferably marginal likelihood/Bayes factors plus posterior predictive diagnostics, or a preregistered information-theoretic alternative. The CBR branch earns evidential credit only for performance not purchased by additional flexibility.
A model can match data and still be weak evidence if many comparably flexible models match equally well. A reviewer therefore asks a comparative question: does this theory predictively compress the observations better than reasonable alternatives?
The benchmark depends on the domain. For low-energy particle physics, the Standard Model plus explicitly needed extensions is the obvious reference. For cosmology, a standard ΛCDM-like baseline and relevant extended models provide a practical comparison. The project does not need to defeat every benchmark numerically on every legacy point; a unification theory may trade a small local fit penalty for major cross-domain compression. But that trade must be quantified rather than asserted.
One clean approach is held-out expected log predictive density. Given independent or appropriately modeled data partitions, compare
Another is a Bayes factor when priors are defensible. Another is minimum description length if the coding rules are fixed before seeing which model wins. The project already values MDL-like economy, but the coding convention must not be selected to flatter the preferred geometry.
Qualitative predictions also matter. If the theory forbids a class of particles, fixes a discrete multiplicity, or predicts a sign that alternatives leave free, successful confirmation can be decisive even if a smooth likelihood comparison is difficult. Conversely, a theory that reproduces only quantities all reasonable models were built to fit gains little discrimination.
Benchmarking also disciplines the meaning of “parameter economy.” If the proposed theory has 13–14 charged reals and the benchmark has 25 under the chosen ledger, that may be meaningful compression; if the proposed theory additionally searched hundreds of structural branches or uses a flexible state function, the comparison must include that cost. The benchmark’s own structural assumptions must be charged under the same coding philosophy.
On held-out data and a common complexity ruler, the branch is at least competitive with domain benchmarks and provides additional cross-domain compression or distinctive predictions that the benchmarks do not obtain for free.
The theory only looks economical because its structural/functional search cost is excluded, or its held-out predictive performance is materially worse without compensating unification evidence.
The championship is not won by showing that our team can score. We must show that the scoring is surprising relative to other teams playing under the same rules.
Championship burden VIII — Independent replay and a genuinely new prediction are the final proof-of-seriousness
Our strongest claim should not depend on our own code or our own interpretation. Someone who was not involved in building the result must be able to reproduce it and test a prediction that was frozen before the answer was known.
At least one independent implementation must reconstruct the decisive result from the mathematical specification and frozen artifacts. For the highest claim level, at least one discriminating observable must be timestamped before target access and later compared once, with repair forbidden until the result is scored.
Even a perfectly documented internal program can share blind spots. The final goalpost is therefore social as well as mathematical: can the decisive claims survive an implementation and review process that was not involved in inventing them?
Independent replay has several levels. At minimum, another agent or researcher should recompute the arithmetic and algebra from the same formulas. Stronger is an independent code implementation from a mathematical specification. Stronger still is an independent reconstruction of the derivation from the frozen source artifacts. The project should distinguish these rather than call all of them “independent validation.”
The reproducer should be given known-answer traps. A competent review must notice the live vacuum issue, the flavor tension, the threshold-provenance nuance, and the difference between gate closure and empirical success. If a reviewer reports “all clear” while missing known defects, that review has failed calibration.
The project should also freeze at least one prediction whose datum is genuinely absent or genuinely withheld. A future experiment is ideal, but a legacy dataset can function as a pseudo-prospective test if access is credibly firewalled and the branch was not historically selected using a close surrogate. The prediction packet should contain the distribution, not merely a point, and state what result would count as failure.
A novel prediction need not be spectacular. A small, precise ratio, sign, hierarchy, spectral feature, or cross-domain relationship can be more convincing than a dramatic but flexible forecast. What matters is that the theory had a real opportunity to be wrong and could not adapt after the fact.
The chronology should be public: commit hash, timestamp, artifact hash, allowed nuisance treatment, and data-release time. If the prediction later succeeds, the evidence is legible without asking the authors to remember what they knew.
Independent implementation reproduces the decisive branch and its uncertainties; the reviewer catches known defects; and at least one locked, genuinely distinguishing prediction succeeds without model revision.
Results depend on undocumented implementation choices, reviewer guidance, or post-release model changes; or prospective predictions fail and are repeatedly replaced without paying the branch-selection cost.
If the theory reaches this point after the previous seven burdens, the rational belief update should be large. The evidence would no longer be primarily that we found a clever representation of known physics. It would be that a rigid structure continued to anticipate nature after we stopped helping it.
Championship burden IX — The theory must make a discriminating prediction, not just an accurate one
A prediction is most persuasive when competing theories would have expected something measurably different. Predicting another number that every reasonable model already predicts does little to distinguish us.
Choose a future or unused observable Y to maximize expected discrimination between the frozen CBR branch and preregistered benchmarks. A natural formal objective is expected log likelihood ratio or Kullback-Leibler separation, subject to experimental feasibility. The prediction must be generated before Y is opened. A success that lies in a region where the benchmark assigned low probability carries far more evidential weight than a generic compatibility check.
Promotion rule
Level-5 language requires at least one prospective discriminating prediction whose value was not used to construct the branch, choose the mechanism, tune a parameter, or set the uncertainty model.
Championship burden X — The result must survive reasonable formulation, scheme, and implementation changes
If the headline result disappears when we change coordinates, numerical solver, renormalization convention, harmless discretization, or another nonphysical choice, we have probably discovered an artifact rather than physics.
Define an equivalence class of admissible formulations \mathfrak F that represent the same physical model. For a claimed observable O, require scheme/implementation variation to remain within the preregistered theoretical uncertainty envelope. Gauge, coordinate, basis, regulator, solver, and discretization changes that should be physically redundant become mandatory invariance controls. Sensitivity to a genuinely physical assumption is allowed—but then that assumption must be exposed as a load-bearing hypothesis.
Championship burden XI — A surprising cross-domain link must survive as one calculation
The most convincing part of a unification claim is when the same upstream structure fixes two things that normally have nothing to do with each other. That is much harder to fake than fitting one domain at a time.
Identify at least one upstream object X whose value/structure was frozen in domain A and propagates through independent maps f_A(X) and f_B(X) into held-out observables in domains A and B. The cross-domain result must retain full provenance and sensitivity accounting. The current R_6 ightarrow M_KK ightarrow t_sep chain is a candidate example of cross-domain inheritance, but its evidential class remains “derived from pre-existing particle-physics calibration,” not zero-parameter prediction.
Championship burden XII — The decisive argument must be explainable without weakening it
If a scientist cannot tell what won the game, what could have made us lose, and which numbers were fitted, the evidence will not be persuasive—even if the appendix is correct.
Every championship-level claim must pass a communication equivalence check: the simple explanation and technical statement must be logically equivalent with respect to scope, assumptions, and falsifiers. The simple layer may omit derivational detail but may not omit a condition that changes the truth value of the claim. Independent readers should be able to recover the claim graph—inputs → derivation → observable → falsifier—from the public-facing text before consulting appendices.
The hostile reviewer’s ten objections—and what would answer each one
The purpose of listing objections in advance is to prevent us from answering only the ones for which we already have attractive material.
- “You chose the geometry because it reproduces the Standard Model.”
Answer required: disclose the candidate search and selection history; identify which observations were constitutive pins; score only held-out consequences as predictions. - “Your compactification scale is just the GUT scale put in by hand.”
Answer required: exact provenance chain, chronology, calibration label, and downstream predictions that did not participate in the particle-scale fit. - “Your ‘few parameters’ claim ignores injected normalizations and branch choices.”
Answer required: effective flexibility ledger and benchmark comparison under the same cost convention. - “Your gate board turns failures into green checkmarks.”
Answer required: separate governance and truth scoreboards; CLOSED-NEGATIVE counts as a physical loss for the mechanism. - “The selected vacuum is a saddle.”
Answer required: full physical Hessian plus controlled quantum correction/metastability calculation. No rhetorical answer substitutes. - “Flavor already falsifies you.”
Answer required: forcedness audit of the one-angle ansatz, same-ruler replay, and prospective repair only if the ansatz is not core-unique. - “The primordial spectrum was inserted through a state choice.”
Answer required: state selection derived from SSBA and parent-induced bulk/boundary dynamics before CMB comparison. - “Dark matter, baryogenesis, and Λ are just missing.”
Answer required: either derive them, openly calibrate them and narrow the claim, or stop calling the current model a complete cosmology. - “With enough mathematical structure you can explain anything after the fact.”
Answer required: frozen candidate grammar, held-out vector, negative controls, branch-search accounting, and novel prediction. - “No independent group has reproduced it.”
Answer required: external replay package and an actual independent report—not another internal agent using the same hidden assumptions.
A theory that can answer all ten does not become infallible. But a reviewer would have far fewer ordinary explanations for why it appears to work.
How much success is enough? Avoiding a new arbitrary goalpost
It would defeat the purpose of this paper to replace vague goalposts with a magic number such as “five predictions” or “three domains.” There is no universal theorem that five successful observables prove a model. The amount of evidence depends on how independent, precise, surprising, and flexibly predicted those observables are.
Therefore the final standard is information-based rather than count-based. The model should produce a held-out predictive distribution whose success is unlikely under relevant alternatives and cannot be cheaply recovered by moving the model. The evidence grows when:
- observables probe independent directions;
- uncertainties are small relative to the theory’s allowed band;
- the prediction is structurally distinctive;
- the model had little flexibility after freeze;
- the target was genuinely unseen;
- the same fixed structure predicts multiple domains;
- reasonable alternatives assign lower predictive probability.
A single sharp sign prediction can therefore be more valuable than fifty fitted curve points. Conversely, a full CMB spectrum can be extremely informative if the primordial state/background and matter ledger were fixed upstream and only a small declared calibration set was used.
The reviewer’s qualitative threshold for “strong evidence” is consequently a conjunction rather than a count: no fatal contradictions + rigid prospective generative chain + successful overconstrained held-out likelihood + benchmark competitiveness + independent replay + at least one high-discrimination prediction.
This standard is deliberately difficult to game. If we later discover a new untested obligation, we may add it because physics demanded it, but we may not redefine an already failed burden so the old result becomes a pass. Version history must show the change and why it was unknowable earlier.
The skeptical-reviewer forced-choice theorem
This is the outcome we should design the evidence package to produce.
After the final test, we do not need every scientist to agree with us. We need disagreement to become concrete. A reviewer should have to say exactly which equation is wrong, which input leaked, which free choice was hidden, which measurement is incompatible, or which replay failed.
Suppose a frozen theory packet passes the twelve burdens, all public artifacts reproduce, benchmark comparisons are predeclared, and a discriminating prospective prediction succeeds. Then a scientifically substantive rejection must attack at least one premise of the evidence chain: mathematical validity, provenance, model-class completeness, uncertainty, data validity, same-ruler mapping, computational replay, or statistical decision rule. This is not a logical proof that the theory is true; it is a proof that generic dismissal is no longer an adequate technical response.
Sector finish line — Geometry and gauge structure
The geometry/gauge sector should not be judged by whether it can be made to resemble the Standard Model. It should be judged by how much of the observed low-energy gauge structure follows once the geometric/category choices and measured anchors are honestly separated.
A reviewer begins by fixing the claim. If the framework says “given the declared CSDR/isometry grammar and the selected internal factors, the low-energy gauge algebra is SU(3)×SU(2)×U(1) with the stated rank and carrier ownership,” then the burden is algebraic and finite. One must compute the isometry/centralizer structure, parities, zero modes, multiplicities, and extra-vector exclusions. A competing carrier that produces unwanted gauge factors is a meaningful negative control. This can be a strong structural result even if the category “forces are internal isometries” is itself a constitutive modeling choice rather than a theorem of nature.
If the stronger claim is “the geometry uniquely explains why nature has the Standard Model gauge group,” the burden becomes much higher. The candidate grammar of geometries and modeling categories must be specified, and rival categories capable of producing the same local interface must be considered. The project’s own gate taxonomy recognizes a category floor here. That is not a failure; it means the correct public claim is conditional on a modeling category. A reviewer will reject wording that silently turns a certified irreducible category posit into a derivation from nothing.
The same distinction applies to the compactification scale. Geometry may determine relations among radii once a scale is fixed, while the absolute scale can still inherit a particle-physics calibration. The reviewer wants a dependency graph showing which statements are topological/exact, which are dimensionless geometric consequences, and which acquire units only after measured anchors or fitted thresholds enter.
The gauge sector becomes convincing when it is both rigid and discriminating. Rigidity means reasonable deformations or alternative embeddings are either excluded by upstream constraints or visibly change predictions. Discrimination means the surviving structure predicts something not used to select it: a forbidden extra U(1), a precise multiplicity, a kernel quotient, a representation relation, a discrete parity pattern, or another nontrivial observable.
Negative controls matter here because group theory is rich enough that many attractive decompositions can be found after the fact. The project should deliberately test neighboring cosets, torus choices, parity assignments, and embeddings that are close enough to be plausible. If many produce the same successful interface, then the result is category-generic rather than Shape-specific. If the frozen Shape survives while close competitors fail for independent reasons, evidence for the chosen structure increases.
We may call the gauge structure a strong prediction of the frozen branch only when the carrier/embedding/parity grammar is frozen independently of the tested gauge observables, the candidate space is sufficiently closed to justify forcedness, and at least one held-out gauge/representation consequence distinguishes the survivor from plausible rivals.
Importantly, this sector does not need to prove cosmology to count as a success. The no-moving-goalposts discipline works both ways: a real gauge theorem should not be discounted because cosmology remains open, just as a gauge theorem should not be promoted into evidence that cosmology is already solved.
Sector finish line — Flavor, masses, and mixing
Flavor is where reconstruction programs are most vulnerable to hidden flexibility because measured masses and mixing matrices provide many numerical targets and many possible parameterizations. The reviewer therefore treats this sector as a stress test of forcedness and parameter accounting.
The first requirement is a clear separation between structural data and continuous fitting data. Structural data include family multiplicity, chirality, allowed Yukawa texture zeros, representation selection rules, discrete symmetries, or geometric overlap classes. Continuous fitting data include normalizations, phases, moduli, and any coefficients adjusted to measured masses or angles. A texture that predicts three families from topology is not weakened merely because mass scales are calibrated; but the calibrated masses do not become independent evidence for the texture.
The second requirement is scale/scheme discipline. Quark masses, Yukawas, CKM elements, neutrino parameters, and phases are not interchangeable raw numbers. The calculation must identify the renormalization scale, running prescription, threshold treatment, basis convention, and uncertainty propagation. Same-ruler comparison is mandatory before calling a residual a sigma discrepancy.
The third requirement is a full-matrix test. A flavor model that uses |Vus| to set its lone angle cannot then advertise |Vus| as predicted. The relevant evidence is the rest of the CKM structure predicted after that calibration, including quantities that are geometrically or algebraically linked. If one element shows a large discrepancy, the reviewer asks whether the relation generating it was uniquely forced or merely one convenient low-dimensional ansatz.
A repair protocol must be prospective. If the one-angle construction fails, the next candidate class should be generated from the building-block logic without access to the offending target. For example, derive all independent phases/invariants permitted by the frozen actor geometry, reduce them by symmetries, and freeze the minimal lawful class. Only then open the withheld CKM element. If the new class contains an extra fitted real calibrated directly to the old discrepancy, the old datum pays for the new parameter and disappears from the evidence score.
Neutrino structure requires the same discipline. If mass splittings, ordering, absolute scale, or PMNS phases are partly measured anchors and partly tuned magnitudes, the ledger must say so. A family-count or chirality result can remain a strong topological prediction while the continuous neutrino sector remains calibrated.
The flavor championship is not “fit every mass.” With enough coefficients, almost any texture can. The championship is compressive overconstraint: a small, structurally motivated parameter set should force several independent ratios, signs, hierarchies, zeros, or mixing relations that survive data.
Promote the flavor sector to “predictive” only after every fitted magnitude/phase is paid, the candidate texture grammar is frozen, the full held-out flavor vector is compared at a common scale/scheme, and no core-forced large discrepancy remains. A failed optional ansatz is retired; a failed unique structural relation falsifies the branch.
Sector finish line — Quantum consistency, state space, and boundary admissibility
Quantum consistency is not one gate. It is the condition that the classical/geometric construction actually defines a lawful quantum theory over the regime in which quantum claims are made.
A reviewer separates several burdens that are easy to blur. Gauge consistency asks whether anomalies cancel and whether gauge fixing/BRST/BV structures are coherent. Positivity/unitarity asks whether physical states have nonnegative norm and evolution preserves probabilities in the relevant domain. Reflection positivity or an equivalent reconstruction condition matters if Euclidean machinery is used to infer Lorentzian quantum theory. UV statements require care: a finite effective theory below a cutoff does not automatically prove a nonperturbative UV completion.
The project’s certified-irreducible terminals can honestly mark an external theorem or unresolved mathematical floor, but those terminals should not be scored as positive empirical evidence. If the framework relies on a named external theorem or constitutive quantum rule, it should be exposed as part of the axiomatic cost.
State Selection and Boundary Admissibility adds a separate burden. An admissible state class is not a selected state. Conditions such as gauge invariance, Hadamard behavior, positivity, finite operational action, regularity, and boundary consistency can remove pathological states while leaving an infinite family. The cosmological prediction is not unique until a state-selection principle—preferably a parent-induced variational or Dirichlet-to-Neumann construction—reduces that freedom to one state or a finite equivalence class.
The reviewer should test uniqueness by perturbing the admissibility rules. If several equally lawful states produce materially different primordial spectra, then the theory contains unresolved state freedom. That freedom must be either derived away or charged like a parameter/function. Calling all such states “admissible” does not make their common predictions unique if they have none.
Boundary terms deserve the same provenance scrutiny as bulk terms. A boundary functional chosen because its Hessian has the desired k-dependence is a cosmological fit. A boundary operator induced by integrating the solved bulk equations with fixed regularity conditions is a derivation. The two can be mathematically identical while carrying completely different evidentiary weight.
Call the quantum/state sector predictive only when the physical state space is consistent, the relevant evolution/observables are well defined, and the boundary state used in predictions is selected by upstream rules rather than target data. Any certified mathematical floor is named as an assumption/dependency, not counted as a derived win.
Sector finish line — Granularity, records, and the first-separability clock
The record/separability calculation is a good example of why the claim ladder matters. It can be a genuine theorem under declared assumptions without being a complete cosmology.
The narrow claim is operational: under the frozen finite-region record protocol, the compactification-controlled energy ceiling and quantum orthogonalization bound imply an earliest time at which a pair of finite records can both be completed and causally packed while their conditional influence lies below the chosen action cell. The result is an existence boundary for a pair, not universal Hilbert-space factorization and not disappearance of entanglement.
A reviewer checks three things. First, the scale provenance: MKK must come from a branch fixed without cosmological target data. The current record supports that chronology, while also requiring the honest label that the absolute scale inherits particle-physics calibration. Second, the theorem conditions: support size, allowed record energy, action-cell quantizer, local influence envelope, and nonlinear/tower/global residuals must be explicit. Third, destructive controls: nearby supports or altered support diameter should move or destroy the crossing in the predicted manner.
The claim should also distinguish a lower bound from an equality. The Margolus–Levitin expression supplies a minimum orthogonalization time at a given energy above the ground state. Saturation by a physical record protocol is an additional assumption unless explicitly constructed. The current paper’s use as an earliest possible record boundary is defensible when labeled accordingly; it becomes overclaim if every physical record is said to complete exactly at that time.
Granularity itself must remain what the project defines it to be: operational distinguishability with a universal positive action/cost floor as a named axiom and measured ℏ as residue/anchor. It does not imply a universal minimum spacetime length. The compactification-derived support length is protocol-specific, not a fundamental pixel size of the universe.
The strongest future evidence from this sector would be cross-domain. If the same independently fixed microscopic scale later controls a novel particle, cosmological, or decoherence feature not used in its calibration, the record theorem becomes part of a larger predictive chain. But until that happens, its scientific value is the narrow theorem itself and the conceptual architecture it enables.
Retain the result as a conditional first-existence theorem when its inputs/protocol are frozen and its residuals are scoped. Promote it toward evidence for cosmology only if the same boundary participates in downstream predictions that survive without cosmological target loading.
Sector finish line — Primordial correlations and the inflation-replacement claim
This sector has already taught the project an important lesson: a mechanism can be intuitively attractive and mathematically wrong for the target. The local-correlation branch lost because its infrared spectral class is wrong. That loss should remain permanent.
The next primordial claim must therefore be stronger than “pre-separation interdependence could correlate distant regions.” Large-scale correlation by itself is not enough. The theory must generate the amplitude, approximate scale invariance, mild red tilt, phase/coherence structure, acceptable tensor content, low non-Gaussianity, and acceptable isocurvature from one state/background calculation.
The correct generative chain is:
Every arrow is load-bearing. A state covariance may mathematically be chosen to give PR(k) proportional to kns−4; that proves only existence. The theory predicts the spectrum only if the covariance follows from the frozen bulk/boundary problem. Likewise, a background chosen to make the pump field yield the desired Hankel index is a fit unless the background already follows from the matter/geometry dynamics.
The reviewer will specifically look for target leakage through asymptotic assumptions. “Choose the regular vacuum” can be ambiguous when multiple regular/Hadamard states exist. “Minimum action” can conceal a boundary choice. “Scale invariance follows from equal information per log k” is not a derivation unless the equal-information rule itself is derived independently and its corrections produce the observed red tilt prospectively.
It is scientifically acceptable if the theory uses one or two calibrated primordial parameters. Many successful cosmological models do. But then those observables are calibration, and the evidence must come from the rest of the spectrum, tensor/non-Gaussian/isocurvature predictions, and downstream CMB/LSS structure.
Do not claim an inflation-replacement mechanism until the native bulk/state system prospectively generates a viable continuum primordial spectrum and horizon-scale correlations. Exact scale invariance or a merely possible nonlocal kernel is not enough.
If this route fails, that does not automatically falsify the particle-physics branch. It falsifies the claim that the current frozen framework already supplies a native alternative primordial mechanism. Again, the goalpost specifies the scope of the loss before the result arrives.
Sector finish line — Full cosmological history
A full cosmology claim is earned only when one native solution survives the entire history. It cannot be assembled from mutually independent epoch-by-epoch fits.
The minimum background ledger contains radiation, neutrinos, baryons, whatever supplies dark gravitational phenomenology, vacuum/late-acceleration contributions, and any geometric/moduli energy that is non-negligible. For each component the theory must state whether the abundance and equation of state are derived, calibrated, or externally anchored. Conservation/exchange equations must close.
Baryogenesis is not optional for a complete history because baryon density affects BBN and CMB. The theory need not derive the observed asymmetry from zero parameters; a calibrated efficiency can be allowed if openly paid. But a claim of native reconstruction is stronger if the sign and magnitude arise from admitted CP violation, anomaly/topology, and nonequilibrium history.
Dark matter is similarly more than Ωc. The theory must provide perturbation behavior: clustering, sound speed, anisotropic stress, free-streaming or interaction scale, and gravitational coupling. A geometric effective component that mimics the background density but fails structure growth does not solve the dark sector.
The late-time vacuum sector must be treated with the same honesty. If Λ is a measured anchor under the current gate taxonomy, then the absolute value is not a prediction. The theory can still predict consequences conditional on Λ. If it later derives a dynamical vacuum law, that becomes a new theory version and must be tested prospectively against expansion data not used to choose the law.
The thermal-history calculation then solves the reaction/transport equations against the native H(T), with phase transitions and entropy changes included. BBN is a crucial firewall because it constrains expansion and particle content long before recombination. Recombination and the Boltzmann hierarchy then generate the CMB spectra. BAO, lensing, matter power, growth, and age are subsequent consequences of the same H(a) and perturbation system.
A reviewer will not require every nuisance astrophysical parameter to come from fundamental geometry. Foregrounds, calibration uncertainties, and instrument response can be marginalized as nuisance parameters. What cannot be marginalized as “nuisance” is a fundamental function introduced to make the cosmic history work.
Call the framework a native cosmology only when one frozen background + matter ledger + state produces a jointly acceptable BBN/CMB/BAO/growth/age history after removing declared calibration observables. Until then, use narrower language such as “cosmological reconstruction program” or “conditional cosmology.”
Sector finish line — What would make “unification” more than a common language?
A theory can unify notation without unifying physics. The reviewer asks whether the same small set of structures creates constraints that propagate across domains in ways that would be unlikely if the domains were modeled independently.
True cross-domain load-bearing structure has three signatures. First, removing it should break multiple predictions at once. Second, the same object should appear through explicit matching maps rather than merely through a shared name. Third, it should create cross-domain relations that were not individually fitted.
For example, if a compactification radius determined from particle unification controls both a KK spectrum and an independently tested early-record boundary, that is a real cross-domain relation—conditional on the scale’s calibration provenance. If the same geometry also fixes a state-selection boundary or feature in a cosmological spectrum, the chain becomes stronger. By contrast, saying “Shape influences everything” without a quantitative dependency is not unification evidence.
Ablation is the strongest test. Remove or alter the 13D Shape while preserving enough ordinary 4D freedom to re-fit local observables. Does the distinctive cross-domain structure disappear? Remove Granularity: does the record/separability theorem or state-selection criterion lose its forcing? Remove SSBA: does the cosmology regain arbitrary state freedom? If an advertised building block can be removed without reducing held-out predictive performance, it is not load-bearing for that evidence.
The reviewer should also compare against modular alternatives. Suppose a conventional GUT plus an unrelated inflationary cosmology fits the same data with similar effective complexity. The unified framework gains evidence only if it reduces total description length, predicts cross-domain relationships, or succeeds on a distinctive held-out observable. Aesthetic unity alone does not beat modularity.
Conversely, the theory need not explain consciousness, every quantum foundation question, or all mathematics before it can count as a unification candidate in physics. Scope should be ambitious but finite: the claimed domain must be enumerated, and the evidence should be judged against that domain.
Use “unification” as an evidentiary claim only when shared frozen structures are quantitatively load-bearing across at least two traditionally independent domains, survive ablation, and improve predictive compression relative to modular alternatives. Otherwise use “common framework” or “reconstruction architecture.”
How I would update belief as the tournament progresses
The purpose of the goalposts is ultimately epistemic: what evidence should rationally make us believe more or less? We do not need a fake precise prior probability, but we do need asymmetric updates.
Some results should move belief only slightly. Recovering a structure that was used to choose the geometry is mostly a consistency check. Reproducing a calibration datum says the implementation can fit what it was told to fit. Showing that a mechanism is mathematically possible is a modest positive update because it removes an impossibility objection but does not show nature selected it.
Other results should move belief strongly downward. A core-forced precision contradiction after same-ruler replay is direct evidence against the branch. A physical vacuum instability is especially serious because it undercuts the background assumed by many downstream successes. A finding that the model’s impressive economy disappears when hidden functional freedom is counted would sharply weaken the unification case even if local fits remain.
The strongest upward updates come from rigidity plus surprise. A locked prediction that was not available during model development, especially one that distinguishes the theory from competitors, has high evidentiary weight. So does a cross-domain relation in which a quantity fixed in one domain predicts a structurally different observable elsewhere. Independent replication multiplies confidence because it attacks shared implementation error.
In Bayesian language, evidence is the likelihood ratio between the proposed theory and alternatives. We need not compute a single global Bayes factor to use the principle:
A result everyone expected under many models has a likelihood ratio near one and should not make us much more confident. A result that the frozen theory assigned high probability and reasonable alternatives assigned low probability can make us much more confident. A result the theory assigned tiny probability should reduce confidence even if our governance taxonomy handles the failure perfectly.
That principle is a powerful antidote to self-deception because it separates methodological virtue from physical likelihood. Reporting a falsifier honestly increases confidence in the method while decreasing confidence in the contradicted mechanism. Both updates can be true simultaneously.
Do not ask “is this result good for the project?” Ask “was this result more likely if the theory is right than if a reasonable alternative is right?” Score the answer, not the emotional valence.
The public-language goalposts
The final anti-goalpost device is linguistic. Public wording should be mechanically tied to the highest cleared claim level.
| Highest level cleared | Allowed language | Language to avoid |
|---|---|---|
| Method only | “A rigorous constraint-based research method that has generated testable structures and falsified several internal hypotheses.” | “Theory of nature,” “unified theory established.” |
| Viable branch | “A mathematically and empirically viable candidate within the tested sector.” | “Explains nature” if prediction/overconstraint not established. |
| Predictive theory | “A predictive model with successful held-out tests after explicit calibration.” | “Parameter-free” unless literally true. |
| Strong unification candidate | “One branch jointly predicts multiple domains with a common parameter ledger.” | “Complete theory” if named domains remain external. |
| Strong evidence | “The accumulated prospective and independently reproduced evidence supports the hypothesis that the framework captures real underlying physical structure.” | “Proven true.” Physical theories remain scope-bound and revisable. |
This table matters because rhetoric can move the psychological goalposts even when the technical ledger does not. If a paper headline says “resolved” and the reader hears “predicted,” the project can mislead without any false equation. The language should do the same accounting as the mathematics.
How repairs are allowed without turning falsification into an impossible game
A good theory-development program must be allowed to learn from failure. Otherwise every early mistake would permanently kill the entire research direction. But repair must be versioned so that learning is not confused with prior prediction.
When a test fails, the first action is classification: did it falsify an auxiliary ansatz, an Actor realization, a building block, or an immutable core statement? The classification must be based on documents that predate the failure where possible. If we redefine what was “core” only after seeing the result, the historical claim must be corrected.
A repair may then be generated from upstream constraints. The crucial rule is that the failed target becomes unavailable during repair selection. The repair agent can know the type of defect—e.g., “the flavor candidate class was incomplete”—but should not optimize against the numerical target that failed. Once a new lawful candidate is frozen, the old failed datum can be used as calibration if needed, but then it no longer counts as evidence for the new version.
The project should maintain theory versions T₀,T₁,T₂,… with explicit diffs: new parameter, new building block, new branch rule, removed claim, and changed prediction set. Evidence belongs to the version that made the prediction. A successful T₂ does not retroactively make T₀’s failed prediction successful.
This is not punitive. It is how scientific progress becomes legible. A research method that quickly converts failures into better constrained theories can be excellent even if early versions are wrong. The eventual theory becomes convincing when the repair rate slows because the frozen structure begins surviving new tests without change.
The strongest sign that the goalposts are finally in the right place will be a long interval in which new data change parameter posteriors and uncertainty bands but do not require new structural building blocks, new branch grammars, or reclassification of old predictions.
Formal same-ruler audit — proving that theory and experiment are actually the same object
Many apparent triumphs and apparent falsifications disappear when the compared quantities are not actually identical. Because this project spans dimensions, effective theories, renormalization scales, boundary observables, and multiple geometric representations, a same-ruler audit must be a formal object rather than a sentence in the discussion.
For every empirical comparison, define a theory-side object XT in its native description and a measured object XE. The comparison is lawful only if there is a published map
such that both sides land in the same comparison space Xcmp. The maps may include dimensional reduction, field redefinition, normalization, renormalization-group running, detector response, basis rotation, projection to gauge-invariant observables, or conversion from a theoretical correlation function to the statistic actually reported by an experiment.
The audit should answer at least seven questions. Dimension: are both quantities defined in the same effective dimension? Frame: Einstein/Jordan/comoving/physical or other frame? Scale: are running parameters compared at the same μ or transferred with a specified RGE? Scheme: MS-bar, on-shell, lattice, or another renormalization convention? Normalization: are fields, generators, mixing matrices, and Fourier transforms normalized identically? Projection: is the theory predicting a gauge-dependent component while the measurement reports a gauge-invariant combination? Statistical object: point estimate, posterior mean, profile likelihood, upper limit, or correlated data vector?
The maps themselves carry uncertainty. If y=R(x,η) depends on nuisance quantities η, the theory-side covariance receives propagated terms
A significance statement that ignores the mapping uncertainty can be too harsh or too flattering. Conversely, declaring a large discrepancy “possibly due to scheme” without quantifying the scheme sensitivity is not an uncertainty analysis; it is an escape hatch.
Same-ruler audits are especially important for quantities like quark masses, Yukawas, CKM/PMNS parameters, threshold corrections, cosmological power spectra, and vacuum terms. A compactification-scale parameter is not directly a collider observable; a primordial curvature spectrum is not directly a CMB Cℓ; a microscopic correlation function is not directly an operational record. Each requires a transfer map.
The reviewer should insist that the map be frozen before scoring the discrepancy whenever multiple plausible conventions exist. If the convention is chosen after looking at which one makes the theory closest to data, that convention choice becomes part of the fit and must be charged.
There is a useful adversarial test: deliberately evaluate the comparison under adjacent reasonable rulers. The correct theory should not require an exceptional convention unless the geometry/dynamics uniquely selects it. If the result swings from 0.5σ to 10σ under conventions that the theory leaves free, then the prediction is not yet sharply defined.
Another test is invertibility at the claimed scope. If the measurement constrains only a projection of a larger theory object, the comparison cannot establish the full object. For example, matching one effective coupling combination does not validate every higher-dimensional component that projects to it. The paper should say “the projected observable agrees,” not “the underlying tensor is confirmed.”
For every decisive empirical result, publish the theory object, measurement object, common comparison space, transfer maps, scale/scheme/frame/normalization, propagated covariance, and adjacent-ruler sensitivity. Only the object that survives this map is scored.
This formalism protects the theory from unfair falsification and protects us from false confirmation. It is therefore neutral but essential officiating.
Uncertainty, truncation, and scheme dependence — a prediction is a distribution, not a favorite number
A reviewer will not accept a precision claim without a complete uncertainty budget. This is particularly important in a reconstruction program where exact topological results, numerical eigenproblems, EFT truncations, threshold sums, measured anchors, and uncomputed higher-order effects coexist.
The first discipline is to type uncertainty by origin. Measurement uncertainty enters through imported anchors. Numerical uncertainty comes from discretization, convergence tolerance, finite precision, or stochastic algorithms. Truncation uncertainty comes from neglected EFT operators, finite KK towers, perturbative order, or asymptotic expansions. Model-form uncertainty arises when multiple admissible realizations remain. Scheme uncertainty appears when a quantity depends on an incompletely fixed renormalization prescription. These categories should not be collapsed into one convenient error bar because they behave differently under future refinement.
If an observable O depends on inputs θ with covariance Σθ, linear propagation gives
but nonlinear dependence may require Monte Carlo or interval propagation. The important point is chronology: the uncertainty model must be specified before seeing whether a wider error bar would rescue the prediction. Retrofitting theory error until χ² becomes acceptable is another form of tuning.
For perturbative series, a practical estimate can use the size and pattern of omitted orders, but the method must be fixed across successes and failures. If two-loop terms are considered negligible for observables that agree and potentially huge for observables that disagree, the uncertainty policy is asymmetric. The reviewer will notice.
KK/tower calculations need a convergence certificate. If a finite truncation N is used, publish the N-dependence and a bound or empirical convergence model for the remainder. A result that stabilizes numerically over N but lacks a uniform analytic theorem can still be useful if the tested domain is finite and the residual is quantified; it should not be promoted to an infinite-tower theorem.
Scheme dependence is a special case. Physical observables should be scheme-independent after all orders, but truncated calculations are not. Varying the renormalization scale or subtraction scheme within a preregistered range is a legitimate sensitivity test. However, if the sign of a load-bearing physical Hessian or threshold correction remains undetermined across reasonable schemes, the theory has not made a sharp prediction at that order.
Discrete ambiguity is not a Gaussian error bar. If two orientations, parities, or branches are both lawful and yield opposite signs, the predictive object is a mixture or a branch class until another rule selects one. Averaging them into a central value can conceal that the theory has not chosen.
The reviewer also asks whether experimental systematics are treated symmetrically. A 14σ nominal discrepancy may be less if shared theory/measurement covariance was omitted; but it remains serious if the full covariance still gives a huge pull. The correct response is calculation, not intuition.
Every scored prediction carries a preregistered uncertainty decomposition, correlations, truncation control, numerical convergence evidence, and scheme sensitivity. The pass/fail verdict is computed from that distribution without widening or narrowing it after target inspection.
A theory earns credibility when its uncertainty bars are predictive too: repeated held-out data should fall inside them at roughly the advertised frequency. Calibration of uncertainty is itself a test.
Candidate-search accounting — how to pay for ideas tried before the winner appeared
A central challenge for any intensive theory-development program is that human and AI creativity can search enormous spaces. Even with zero continuous parameters, trying many structures against known data can create apparent miracles. The final evidence standard must therefore account for model search.
There are several kinds of search. Explicit enumeration tries a known finite list of geometries, representations, or kernels. Sequential repair modifies a building block after a failure. Prompt-level search asks multiple agents or conversations for candidate mechanisms. Implicit search occurs when a researcher uses domain knowledge to invent one candidate while already knowing the target. Only the first is easy to count, but all affect evidentiary interpretation.
The project does not need to assign a precise “number of thoughts” penalty. It does need to distinguish discovery evidence from post-freeze prediction evidence. Results used during search are primarily selection data. Once a branch is frozen, new held-out results become prediction data. This simple chronology removes much of the need for speculative search penalties.
For legacy observables that inevitably guided development, a model-comparison analysis can still charge branch multiplicity. If a finite set B of serious candidates was searched, the model evidence marginalizes over them rather than conditioning on the winner. If the set is open-ended, minimum-description-length or explicit coding of structural choices can provide a conservative complexity estimate.
Sequential building-block evolution should be versioned. Suppose a gate failure reveals a missing commutator-closure constraint. Adding that reusable constraint may be a genuine methodological advance. But the gate that revealed it becomes training data for the new block. Evidence for the new block comes from later gates/problems that close without further target-specific repair. This is exactly analogous to training and test sets in machine learning.
The same principle applies to the new State Selection and Boundary Admissibility block. Its existence was motivated by a cosmology failure: the frozen theory did not select a unique global state kernel. Therefore the failure that motivated SSBA cannot also be counted as independent evidence that SSBA predicted the right kernel. The next independent application is the real test of whether SSBA is a reusable physical principle rather than a bespoke repair.
A reviewer will also ask whether failures are archived. If only successful candidate documents survive, the historical search appears much narrower than it was. The failure registry is therefore evidence discipline, not embarrassment. It allows future model-selection analyses to understand how much flexibility the process exercised.
Publish the serious candidate/repair history, identify which observations were visible during each choice, treat discovery data as selection/calibration rather than prediction, and reserve strongest evidence for post-freeze applications that the new structure did not see.
This goalpost does not punish creativity. It simply locates where creativity ends and prediction begins.
The prospective prediction protocol — exactly how we should run the next championship test
A final reviewer document should not merely say “make predictions.” It should specify the operational sequence so a future success cannot be retrospectively reinterpreted.
Step 1 — Choose the claim before the target
Write one sentence describing the prediction and its scope. Example structure: “Given frozen theory version T and calibration set C, observable vector Y in domain D is predicted to follow distribution P with no further structural adjustment.” The claim includes what will not be predicted.
Step 2 — Freeze the candidate grammar
List admissible branches, discrete signs, operator bases, truncation rules, and the algorithm that selects among them. If the grammar is not closed, label the test exploratory rather than confirmatory.
Step 3 — Freeze inputs and calibration
Publish all measured anchors and fitted parameters. Each fitted parameter is paired with the calibration datum or likelihood block that pays for it. The calibration data are tagged so they cannot later appear in the prediction score.
Step 4 — Freeze the computation
Hash code, solver settings, random seeds where relevant, convergence tolerances, numerical libraries, and the output schema. For stochastic or posterior predictions, specify the distribution and sampling procedure.
Step 5 — Freeze the same-ruler map
State exactly how theory output becomes the measured statistic, including scales, units, frame, projection, selection effects, and nuisance marginalization.
Step 6 — Freeze uncertainty and falsifiers
State the predicted interval/distribution and the decision rule. The falsifier should be finite and decidable. If multiple outcomes correspond to local vs branch-fatal losses, state the classification rules.
Step 7 — Timestamp the packet
Use a public repository commit, notarized hash, preprint timestamp, or another independent archive. The critical property is that later readers can verify the packet predated target access.
Step 8 — Open the target once
Run the frozen comparison without repair. Record the result even if it fails. If a software bug is discovered, document the bug and rerun under a versioned correction; do not change physics choices in the same patch.
Step 9 — Score against benchmarks
Compute the same held-out likelihood or qualitative discrimination for relevant alternatives. Use the same data and uncertainty treatment.
Step 10 — Only then begin repair
If the theory loses, archive the verdict. A new theory version may be developed, but the failed target becomes part of the training history and cannot be recycled as a clean prediction.
If the next major prediction succeeds under this sequence, we will not need to argue about whether we subconsciously tuned it. The chronology itself will answer the objection.
The joint cross-domain score — one scoreboard for the whole theory
The final temptation to move goalposts appears when each domain is scored independently. A branch may accumulate many local passes while hiding that different sectors use incompatible parameters, or that one severe loss is being diluted by many easy wins. The theory therefore needs a joint scoreboard with veto conditions.
Let D1,…,Dm denote domain-level held-out data blocks—precision particle physics, flavor, record/quantum tests, BBN, CMB, BAO/LSS, and prospective novel data. If the conditional independence structure is known, a joint predictive likelihood factors accordingly; otherwise the full covariance/hierarchical model should be used:
The exact scalar score is less important than preserving three rules. No refitting between domains: shared parameters have one posterior/distribution. No averaging away fatal contradictions: a branch-level veto is applied before summing log scores. No double counting: correlated measurements or derived summaries of the same data are not treated as independent wins.
The scoreboard can therefore have two layers. Layer A contains veto tests: physical stability, anomaly/unitarity consistency, and any confirmed core-forced empirical contradiction. Failure of a veto rejects the frozen branch regardless of how well other curves fit. Layer B contains comparative predictive evidence for branches that survive Layer A.
This resembles sports playoffs better than a season point total. A team cannot compensate for fielding an ineligible player by scoring extra goals; eligibility is a prerequisite. Likewise, an unstable physical background is not a −10 point penalty that can be offset by ten minor numerical matches.
For surviving branches, report domain-by-domain predictive scores as well as the joint score. This prevents a good average from hiding systematic weakness in one domain. A theory that is slightly worse than benchmarks in several legacy domains but dramatically more compressive across all of them may still be interesting; the trade must be explicit.
The joint scoreboard should include model complexity only once. Shared structure that predicts multiple domains earns compression because the same parameter is not repaid separately. Domain-specific free functions are charged in their domains and weaken the unification evidence.
Prediction correlations created by shared upstream anchors should be propagated. If uncertainty in MU shifts both a particle threshold and an early-universe clock, those predictions are correlated and should not be counted as two fully independent pieces of evidence. On the other hand, a successful cross-domain relation can be powerful precisely because the correlation itself is predicted.
No veto failure; one shared branch/parameter ledger; no double counting; acceptable domain-level fit; and competitive joint held-out predictive performance after complexity. This, rather than the count of resolved gates, is the true theory scoreboard.
A simulated final referee report: what I would write if the manuscript arrived today
It is useful to state the likely review in plain language, because formalism can obscure the actual evidentiary position.
Strengths. The program is unusually explicit about assumptions, provenance, falsifiers, and negative controls. The gate methodology has demonstrated the ability to find and preserve failures instead of smoothing them away. Several structural calculations are nontrivial and reproducible in principle. The record/separability work is carefully scoped and the local primordial-correlation no-go is a genuine negative result. The recent State Selection and Boundary Admissibility block addresses a real missing logical step rather than simply inserting a cosmological covariance.
Primary concern 1: physical vacuum. The selected Shape-doublet direction is presently recorded as a live falsifier/saddle with no frozen-data stabilization certificate. Because the same selected background underlies much downstream calculation, this is not a minor appendix issue. The manuscript should not claim a physically realized unified branch until the physical Hessian/metastability problem is resolved.
Primary concern 2: precision flavor. The ~14σ |Vtd| discrepancy must be adjudicated at the level of forcedness. If the one-angle relation is a uniquely forced consequence of the frozen core, this is strong evidence against the branch. If it is an auxiliary ansatz, the manuscript should retract the theory-level forcedness claim, retire the ansatz, and test a prospectively derived replacement.
Primary concern 3: parameter/search accounting. The project has corrected earlier optimistic economy claims and now charges substantially more injected reals. This correction is a strength, but the final evidence assessment must also include discrete branch search and any functional freedom introduced by state/cosmology sectors.
Primary concern 4: cosmological generative closure. The current framework does not yet supply a complete native background, matter ledger, and uniquely selected primordial state. Downstream cosmological transfer calculations using imported parameters are useful controls, not evidence that the cosmology has been derived.
Primary concern 5: independent evidence. Much of the work is internally generated and reviewed. The next evidentiary step should prioritize independent replay and a locked prediction rather than additional internally compatible reconstructions.
Recommendation today. Continue the research program; do not publish the strongest “theory is correct” interpretation yet. Publish the method, structural results, negative results, and precise open/falsifier ledger. Treat the vacuum and flavor questions as decision gates. If those clear, execute the native background/state calculation under a predeclared prediction protocol. A successful cross-domain held-out package would warrant a major reassessment.
What would change the recommendation. A frozen-data vacuum stability certificate, prospective resolution of the flavor forcing chain, native state/background derivation, and successful independent held-out prediction would move the manuscript from “interesting rigorous program” to “serious predictive unification candidate.” Independent replication plus a novel successful prediction would justify substantially stronger language.
The hard-stop constitution — results that end the current branch rather than create another appendix
The most important way to keep goalposts fixed is to define, before the next calculation, which outcomes are genuine stop conditions. Otherwise every failure can be redescribed as “one more thing to solve,” and the theory becomes impossible to falsify in practice.
A hard stop is not triggered by every failed submodel. It is triggered when the failure reaches an object the current branch has declared immutable or uniquely forced. The burden is therefore two-part: first establish the logical ownership of the prediction, then score the evidence.
Hard stop H1 — no physical background
If the selected compactification/background has a reproducible physical instability and no stabilization follows from already admitted dynamics under a preregistered completion, the current branch is rejected. A later theory version may add new physics, but it is a different branch and the prior downstream predictions become conditional explorations on an unrealized background.
Hard stop H2 — core-forced precision contradiction
If an observable is proven invariant across the complete admissible candidate grammar and disagrees decisively with a reliable same-ruler measurement, the branch is rejected. The discrepancy need not be exactly “5σ”; the operative standard is decisive predictive incompatibility after theory and measurement covariance are included. A ~14σ result, if truly core-forced and independently replayed, is well beyond the point where wording changes are an adequate response.
Hard stop H3 — mathematical inconsistency in the claimed quantum regime
An uncancelled gauge anomaly, physical negative norm, non-unitary evolution where unitarity is claimed, or contradiction in the constraint algebra is fatal to that claimed regime unless an already-declared completion supplies the missing term. A certified external theorem dependency is not the same as an inconsistency; the distinction must remain explicit.
Hard stop H4 — prediction requires unrestricted target-dependent functional freedom
If a flagship prediction can only be recovered by introducing an arbitrary function whose form is selected from the target data, the native-prediction claim is rejected. The resulting model may remain a phenomenological fit, but the explanatory mechanism has failed its stated burden.
Hard stop H5 — cross-domain incompatibility
If the particle sector and cosmology require mutually incompatible values of a supposedly shared frozen quantity, or require incompatible discrete branches with no lawful superselection interpretation, the unification claim fails. One can retain the separate sector models, but not the assertion that one branch explains both.
Hard stop H6 — decisive prospective failure after full freeze
If a high-discrimination prospective prediction is locked under this constitution and later fails decisively, the exact claimed mechanism loses. Repeatedly issuing new predictions from repaired versions is legitimate research, but the failed version’s evidence cannot be erased. A mature theory should eventually survive without structural repair.
Some outcomes are explicitly not hard stops. A fitted parameter is not fatal if openly paid. Failure to derive a measured anchor is not fatal if the claim is conditional on that anchor. An optional ansatz can fail without killing the core if optionality was documented before failure. An external mathematical wall can limit scope without making the model internally inconsistent.
This constitution is intentionally uncomfortable. Its value is that future success becomes more meaningful because failure was genuinely possible.
The chronology standard — proving that a prediction existed before the answer
For this project, provenance is not merely “where did the number come from?” It is also “when did each choice become fixed relative to the data it later explains?” Chronology is the strongest defense against unconscious circularity.
The reviewer should reconstruct a time-directed graph. Every node is a document version, parameter value, branch choice, building block, or prediction. Every edge says “this later object was allowed to depend on that earlier object.” Data-release nodes are included. A valid prospective prediction has no directed path from the target datum to the theory choices that generated the prediction.
This is stronger than merely saying “we did not consciously fit it.” If a close surrogate observable was used to select the branch, the target may not be fully independent even if its exact number was hidden. The paper should therefore list proxy leakage: observables or qualitative facts that strongly constrain the same theory direction.
There are several evidence grades:
- Grade C0 — constitutive: the datum directly defines or pins the category/anchor. It carries no predictive score.
- Grade C1 — calibrated: the datum determines a continuous/discrete parameter. It carries no predictive score for that parameterized version.
- Grade C2 — historically visible but withheld in current execution: useful as a robustness test, but weaker because human model design may already encode it.
- Grade C3 — genuinely firewalled legacy datum: neither builders nor selection agents access the datum or close proxies during branch construction.
- Grade C4 — prospective datum: prediction packet predates the measurement/data release. Highest evidentiary grade.
The compactification-scale/cosmology example illustrates why this matters. The absolute scale appears in particle-physics work before the early-universe papers, which protects the cosmology result from direct cosmological target tuning. But the scale’s particle-side calibration provenance remains part of the claim. Chronology prevents one kind of circularity; parameter classification prevents another.
Building blocks also need chronology. If a new block is invented because Gate X failed, then Gate X is training evidence for the block. The next unrelated gate is the block’s first clean validation opportunity. This is especially important for SSBA: it was motivated by failure to derive a unique primordial kernel, so cosmological state selection cannot itself be counted as an independent proof that SSBA was “predicted from first principles.” Its value will rise when it solves a separate boundary/state problem without target-specific modification.
Git history or timestamped archives are ideal because they make the chronology externally inspectable. If only local file timestamps exist, they are weaker evidence. The final publication packet should include commit hashes for every decisive pre-target freeze. If a historical commit does not exist, the paper should say so rather than fabricate certainty from later reconstructed provenance.
A reviewer may also request a blind replay. Give a new agent the pre-target artifacts only, exclude the target and downstream documents, and ask it to regenerate the prediction. If it cannot, the historical chain may depend on undocumented context. If it can, confidence in the provenance increases.
For every flagship prediction, publish the dependency graph, target/proxy visibility grade, freeze timestamp/hash, and a replay using only pre-target artifacts. Strongest claims are reserved for C3/C4 evidence.
The master reviewer matrix — claim, proof, falsifier, repair, and promotion
This matrix is the practical endpoint of the document. It is the place we should look before saying “we crossed the goalpost.” Each major claim has one proof burden, one falsifier class, one allowed repair rule, and one promotion condition.
| Claim | Required proof/evidence | Decisive falsifier | Allowed repair | Promotion when... |
|---|---|---|---|---|
| Selected Shape is physically viable | Stationary solution; physical Hessian; loop/metastability control | Persistent physical unstable mode on frozen branch | Only prospective already-admitted completion; otherwise new theory version | Stability certificate independently replayed |
| Gauge structure is forced within grammar | Closed carrier/embedding/parity classification + negative controls | Lawful rival in same grammar gives same interface with different claimed “forced” structure | Expand grammar and downgrade previous forcedness | Held-out gauge/representation consequence succeeds |
| Flavor mechanism predicts CKM/PMNS structure | Paid parameter ledger; same-scale full-matrix prediction | Core-unique high-significance mismatch | Retire optional ansatz or new prospectively derived theory version | Held-out flavor vector survives without target-specific re-fit |
| Compactification clock is cross-domain evidence | Pre-cosmology scale provenance + exact record theorem | Cosmological input discovered upstream or theorem conditions fail | Relabel as calibrated/conditional if provenance narrower | Additional downstream feature tied to same scale succeeds |
| Granularity is load-bearing | Ablation changes record/state predictions in preregistered manner | Removing cell law leaves all claimed predictions unchanged | Narrow role; do not call it explanatory for unaffected outputs | Independent applications require same rule |
| SSBA selects physical state | Admissibility + unique/finite state selection from upstream principle | Large lawful state family remains with target-sensitive predictions | Derive additional upstream constraint; paid if new | Independent boundary problem solved without modification |
| Primordial mechanism replaces local/inflationary explanation | Native saddle + induced kernel + viable scalar/tensor/non-Gaussian spectrum | Derived spectrum wrong class or requires target-chosen kernel | New theory version; local closed-negative branch remains dead | Prospective spectrum features survive data |
| Baryogenesis is native | Concrete CP/out-of-equilibrium mechanism + predicted ηB distribution | No asymmetry or wrong sign/magnitude under frozen dynamics | Calibrate efficiency and narrow claim, or add new version | Additional correlated baryon/lepton consequence succeeds |
| Dark sector is explained | Derived stress-energy + perturbation/clustering dynamics | Background match but structure/lensing/growth failure | Calibrated abundance allowed; arbitrary target-shaped function not native | One dark mechanism fits independent gravitational probes |
| Native cosmology | One H(a)+state+matter ledger through BBN/CMB/BAO/growth/age | Needs incompatible histories or target-specific free functions | New version; calibration paid and claim narrowed | Joint held-out cosmology competitive with benchmark |
| Unification claim | One branch, one ledger, load-bearing cross-domain relations | Domains require incompatible branches/values | Downgrade to modular/common-framework claim | Ablation + cross-domain prediction demonstrate shared structure |
| Strong evidence theory captures nature | All vetoes clear + held-out overconstraint + benchmark competitiveness + independent prospective success | Independent replay fails or novel prediction decisively fails | New theory version; evidence resets for changed structure | External reproducer and new data both agree |
The repair column is as important as the falsifier column. It prevents two opposite mistakes: killing the entire research program because an auxiliary construction failed, and protecting the current frozen branch by declaring every failure auxiliary after the fact.
The promotion column also fixes the recurrent “we are done” problem. Clearing the proof burden establishes the scoped claim in column one. Promotion requires a new kind of evidence. For example, deriving a lawful state does not automatically promote to “correct primordial cosmology”; the state’s spectrum must survive data. Deriving H(a) does not automatically promote to “complete universe”; the same history must survive BBN, CMB, BAO, growth, and age.
Before every future status change, identify the row, show the exact evidence that clears the current proof burden, and identify the next promotion condition. If the next condition has not been met, stop at the current claim level.
Tournament rules v2 — a story humans can follow without lowering the physics bar
The sports narrative is not decoration. It gives the reader a memory structure for a complicated proof burden.
| Stage | Question | What counts as advancing? | What sends the branch home? |
|---|---|---|---|
| Eligibility check | Is this even a lawful physical theory version? | Mathematical consistency, physical vacuum/background, provenance. | Hard quantum inconsistency or nonviable background. |
| Regular season | Can it reconstruct known sectors honestly? | Same-ruler fits/predictions with complete parameter accounting. | Core-forced precision contradiction. |
| Playoffs | Can it predict held-out data after calibration? | Overconstrained, cross-domain held-out success. | Target-dependent repair or incompatible branches. |
| Conference final | Does cosmology emerge natively? | Background + state + matter ledger → BBN/CMB/BAO/growth/age. | Imported target-shaped state/history or failed history. |
| Championship | Can it beat reasonable alternatives on something genuinely new? | Prospective discriminating prediction plus benchmark comparison. | Prospective failure under frozen rules. |
| Replay review | Can outsiders reproduce the win? | Independent implementation and public artifacts. | Non-reproducibility or unresolved implementation dependence. |
A reader should always know the score and why it matters. A beautiful calculation in the regular season cannot substitute for winning the championship.
The narrative stages correspond to nested evidentiary sets. Advancement requires satisfying all earlier hard constraints; later evidence cannot compensate for an earlier fatal failure. This makes the tournament a monotone decision process rather than a rhetorical points system.
The championship game plan — the exact sequence that maximizes information and minimizes self-deception
Once the goalposts are fixed, the project should choose calculations by expected decision value rather than excitement. The best next calculation is the one most likely to tell us whether the frozen branch deserves to live, not the one most likely to produce another attractive number.
Phase 1 — Eligibility: settle whether the current branch is allowed to take the field
The first two tests are the vacuum and flavor red zones because they can invalidate the present branch before we invest in full cosmology. For vacuum stability, construct the full physical fluctuation operator around the selected compactification, identify gauge/symmetry zero modes, compute the relevant one-loop/Casimir/Coleman-Weinberg contribution with all currently admitted fields, and quantify scheme/orientation ambiguity. The output is not “a plausible stabilizing sign”; it is a spectrum or metastability distribution with a preregistered verdict. If no frozen completion can remove a physical negative mode, reject the current branch.
For flavor, freeze a forcedness audit before touching the discrepant target. Enumerate the exact assumptions that collapse the mixing structure to the one-angle ansatz. Determine which are immutable consequences of Shape/Actors and which are model conveniences. If the one-angle form is not uniquely forced, demote it before repair and construct the next candidate class from upstream invariants while withholding |Vtd|. If it is uniquely forced, replay the same-ruler prediction with full covariance; a persistent decisive mismatch is branch-fatal.
These tests should be performed by separate execution and review agents. The reviewer is prohibited from suggesting the repair. The original and reviewed outputs are both archived. This matters because the psychological pressure to preserve the branch is highest at exactly these two tests.
Phase 2 — Generative closure: make cosmology emerge rather than enter as input
If the branch remains eligible, solve the native 4D effective cosmological problem from the compactified parent action. The deliverable is a background solution Φ̄(a) with its Hamiltonian constraint, conserved/exchange currents, and complete source ledger. Any matter component not derived is explicitly marked calibrated or external. Do not open CMB/BAO targets while choosing the background ansatz beyond facts already designated as constitutive.
Next perform the physical perturbation reduction and apply State Selection and Boundary Admissibility. The key output is the parent-induced boundary/on-shell Hessian or Dirichlet-to-Neumann operator that fixes the Gaussian state kernel—or proves that the theory does not uniquely select one. If a finite state ambiguity remains, publish the discrete class and its predicted spread. If an arbitrary functional ambiguity remains, the native primordial claim is not closed and the theory should not be tuned against ns under the guise of “selection.”
Only then generate the primordial scalar/tensor spectra, non-Gaussian/isocurvature predictions, and any distinctive features. The prediction packet is frozen and hashed before comparison. If one or two scalar parameters must be calibrated, declare the calibration observables in advance and remove them from the evidence score.
Phase 3 — History: force one solution through the universe
Complete the baryon and dark-sector ledgers. Baryogenesis may emerge natively or may require a paid efficiency parameter; both are scientifically legitimate but represent different claim levels. The dark sector must specify perturbation behavior, not just a density. Vacuum/late acceleration must likewise be typed as derived, calibrated, or measured-anchor. The goal is one stress-energy and interaction system that closes the background and perturbation equations.
Run thermal history without epoch-by-epoch retuning. Electroweak/QCD transitions, neutrino decoupling, entropy transfer, and BBN reaction networks should use the same H(T). The predicted abundance vector is compared with data after any declared baryon calibration is removed. Then recombination and the Boltzmann hierarchy generate TT/TE/EE/lensing spectra and the matter transfer function. BAO, growth, weak lensing, and age follow from the same history.
The scoring packet should contain domain-level and joint likelihoods, benchmark comparisons, and residual plots—but also the parameter Jacobian/Fisher structure so a reviewer can see whether many plotted points represent genuinely independent constraints. Calibration swaps and leave-one-domain-out replays are run before any model repair.
Phase 4 — The real final: predict something we did not build around
If the theory survives the legacy evidence, stop modifying the core and choose a high-discrimination prospective target. The best target is one for which the framework makes a relatively sharp prediction while reasonable alternatives allow a broader range or a different qualitative result. It need not be dramatic. A precise relation among particle observables, a sign/hierarchy, a spectral feature tied to the compactification scale, or a cross-domain correlation can be ideal.
Publish the prediction distribution, same-ruler map, nuisance treatment, and fatal threshold before the relevant data are available. Ask an external reviewer to verify the packet before the outcome. When the measurement arrives, score it once. A success at this stage is qualitatively different from a legacy reconstruction because the theory had no opportunity to learn the answer.
Decision checkpoints
Checkpoint A: after Phase 1, either the current branch is physically eligible or it is rejected. Do not proceed to Phase 2 for evidentiary purposes if eligibility fails; any further work is exploration of a replacement version.
Checkpoint B: after Phase 2, either the theory has a native background/state generative system or its cosmology remains phenomenological/conditional. Do not call a tuned kernel a derived primordial theory.
Checkpoint C: after Phase 3, either one branch provides a competitive overconstrained cross-domain history or the unification claim is downgraded. Local successes may remain publishable.
Checkpoint D: after Phase 4 and independent replay, a successful distinctive prediction upgrades the rational belief status from “serious candidate” toward “strong evidence of real underlying structure.” A failure is recorded against the frozen version and triggers a new-version decision, not a moved goalpost.
Evidence replay: put the current branch against the new goalposts
The following rounds reuse the existing decisive tournament, but each is now preceded by the reviewer goalpost it is actually supposed to clear. This prevents a local result from masquerading as a global finish.
The old tournament remains useful evidence. What changes here is the officiating: each result is scored against the claim ladder and the final eight-part championship burden.
The clock audit
Did the 10^-41 s result secretly inherit cosmology through the compactification scale?
Every absolute scale used in a cross-domain prediction must have an upstream provenance chain that predates the target domain. A calibrated particle-physics scale is allowed, but it must be called calibrated and cannot be counted as a zero-parameter prediction.
The microscopic record/separability papers derive the first completed finite record at
at the symmetric chamber center, using MKK=1/R6 and R6=(2πMU)−1. The decisive audit question is not whether the arithmetic is correct; it is where MU came from.
The project record supports a strong but narrower claim than “Shape predicts the scale from nothing.” The scale predates the cosmology work and is part of the particle-physics branch. Therefore the early-universe time was not tuned against cosmological age, CMB tilt, BBN, or BAO. But the current gate authority also contains an explicit internal tension in the threshold provenance: one early summary labels the threshold triple derived, while the detailed SG-7 audit later says the printed column-total magnitudes are injected and are not regenerated by the target-blind first-principles computation.
GATES_SOURCE_OF_TRUTH.md, lines 914; 1257–1265; 15038–15040; 15203–15206Early sections call the threshold vector derived. The detailed SG-7 section says the printed δ-triple is injected, that only per-packet signs are structurally derived in their own normalization, and that the net column totals require an uncomputed inter-structure normalization.
The threshold-vector conflict must be stated, not harmonized
The older GUT Appendix G describes the threshold vector as generated by a finite heat-kernel ledger and says no row is free. The later gate source-of-truth is stricter: it records that a target-blind reconstruction under a canonical untuned normalization returned
rather than the printed
Controlling status
Ruling. The cosmology is not circular with respect to cosmological data, but the absolute time is not presently a zero-parameter geometric prediction. It is a cross-domain consequence of a pre-existing particle-physics calibration. That remains scientifically useful, provided the label is exact.
The free-parameter combine
Can we tune the theory until the universe looks right and then call the result a prediction?
Every continuous scalar, discrete branch choice, normalization, free function, prior, and target-selected rule must be charged. A fit is legitimate; hidden flexibility is not.
The answer must be no. Fitting parameters is legitimate physics; hiding them is not. The strongest version of the cosmology program therefore keeps a parameter constitution with distinct classes: measured root/anchor, frozen construction input, derived output, cosmology-facing fitted parameter, nuisance parameter, and held-out prediction.
The current GUT economy audit is particularly important because it corrects the seductive “four inputs” story. The gate source explicitly charges roughly 13–14 measured/tuned reals in the active branch when sector normalizations, the threshold triple, Higgs modulus, and neutrino-sector magnitudes are counted honestly. That does not invalidate the model; it changes the model-complexity claim.
GATES_SOURCE_OF_TRUTH.md, lines 174; 908–914; 942The source charges four headline anchors plus roughly nine to ten additional injected reals, including the threshold triple. The old ~4× economy headline is retired; the source reports an honest ~1.8× economy exhibit under its chosen bit-cost ledger.
Five parameter classes — never mix them
| Class | Examples | Can be adjusted to cosmology? | Claim treatment |
|---|---|---|---|
| Upstream measured root | ħ, M_Pl, laboratory gauge couplings | No post-cosmology adjustment | Paid anchor |
| Pre-existing particle calibration | M_U/R_6 on the current branch | No | Cross-domain calibrated input |
| Theory-derived object | record criterion, local-kernel no-go, transfer equations | No | Prediction/theorem conditional on upstream inputs |
| Cosmology fit parameter | A_G, ν_G, η_B, ρ_X0, ρ_V0 | Yes, under this Constitution | Calibration, not prediction |
| Instrument/nuisance parameter | foreground amplitudes, beam/calibration terms | Yes within likelihood | Measurement model; not fundamental theory |
Forbidden category
Calibration map: exactly one datum pays for each knob
| Parameter | Calibration datum | Frozen immediately after | Held-out consequences |
|---|---|---|---|
| A_G | A_s at one pivot | scalar amplitude match | shape of CMB spectrum away from pivot; lensing/matter power once transfer is fixed |
| ν_G | n_s at the same pivot | tilt match | running, cutoffs/oscillations, all non-power-law structure |
| η_B | one baryometer (prefer D/H or a predeclared baryon-density datum) | baryon number | Y_p, independent CMB baryon loading, acoustic response |
| ρ_X,0 | one matter/equality datum | dark abundance | growth rate, lensing, detailed peak heights, matter P(k) |
| ρ_V,0 | one low-z acceleration/distance datum | vacuum normalization | rest of H(z), BAO/SN shape, H_0 under closure, cosmic age |
Not allowed
The calibrated cosmology lane is allowed to exist, but it is a different claim from native prediction. The clean rule is one calibration observable per fitted scalar parameter, frozen afterward. A free functional kernel K(k), w(z), or transfer function is not one knob; unless a finite basis and complexity penalty are fixed in advance, it can absorb arbitrarily much data.
After every fit, rerun the scorecard with all calibration observables removed. If the impressive agreement disappears, the model fitted the data rather than predicted it. If a small calibration set predicts many held-out observables, that is real leverage.
The microscopic clock
The claim is modest but clean: given the frozen microscopic scale and the finite-record protocol, there is an earliest time at which two completed finite records can exist independently under the modeled local influence bounds. This does not say every subsystem in the universe factorizes at that time.
The exact claim is an existence statement, t_sep^∃, obtained by intersecting the causal packing condition with the Margolus-Levitin record-completion lower envelope and then proving the modeled massless/heavy influence envelopes lie below the frozen action-cell threshold. Its absolute normalization inherits the pre-existing particle-physics calibration of M_KK and measured ħ; nonlinear/global/infinite-tower residuals remain outside the unconditional theorem.
Does the record/separability calculation actually follow from its stated assumptions?
The first-record/separability result must follow reproducibly from the stated finite-region protocol and its declared inputs. It need not prove all of cosmology, but it must not claim more scope than it actually establishes.
This round tests the narrow theorem, not cosmological completeness. The ingredients are a compactification-controlled energy scale, a finite-region record protocol, the Margolus–Levitin lower bound for orthogonalization, a causal packing condition, and an operational record-cell quantizer. Under the chosen primitive protocol the record bound is the controlling condition:
The most conservative local massless influence envelope used in the existing calculation gives C/ℏ ≤ x/(2x−1), which is already below one cell at x=π/2. Heavy/KK contributions fall faster in the modeled regime. The theorem is explicitly an existence result for a pair of finite records, not a proof that every subsystem factorizes, and not a statement that collapse signals propagate.
Rigor theorem 3 — first-record/separability boundary
The compactification clock is fixed by the frozen radius, not by cosmological data:
With one primitive support diameter of order ℓKK and the Margolus–Levitin orthogonalization bound at the admitted primitive energy ceiling, the first-existence operational-separability branch obeys
This is a conditional theorem about the declared finite-region protocol. It is not a theorem of universal Hilbert-space factorization, and global/topological sectors remain outside the local separability statement.
Rigor theorem 4 — finite-region influence functional
For a lawful split into retained variables q and integrated variables χ, the closed-time-path generating functional defines the influence action by integrating χ on the forward/backward branches. To quadratic order the most general Gaussian form can be written
DR is the reaction/retarded kernel and N the noise kernel. This is the mathematically correct place to derive time-dependent conditional influence. A static KK mass gap alone is not permission to write an exponential decay in coordinate time.
The remaining caveats are real: the nonlinear influence remainder, infinite KK tower uniformity, and genuinely global/shared modes are not fully closed by the local calculation. But those are declared conditions rather than hidden assumptions.
The local-correlation quarterfinal
This round is a loss, and that is useful. Ordinary local correlations that die with distance cannot produce the nearly scale-invariant primordial pattern seen across the sky. We therefore cannot sell the local separability mechanism as a replacement for the cosmological horizon mechanism.
For an integrable finite-correlation-length kernel, analyticity at k=0 gives P(k)=P_0+O(k²), hence dimensionless power Δ²∝k³ in the infrared. For C(r)∝r^{-α}, 0<α<3, Fourier scaling gives P(k)∝k^{α-3} and therefore n_s=1+α>1. Both classes are incompatible with a slightly red near-scale-invariant scalar spectrum.
Can the microscopic local interdependence mechanism grow into the observed primordial spectrum?
A proposed mechanism for primordial correlations must produce the required infrared structure from its own equations. A mechanism that mathematically yields the wrong spectral class is eliminated, not retained as an intuition.
This is where the theory took a clean loss, and the loss should count. For an integrable finite-correlation-length real-space kernel C(r), the Fourier spectrum is analytic at k=0:
Likewise, a decaying power-law tail C(r)∝r−α with 0<α<3 gives P(k)∝kα−3 and therefore ns=1+α>1. The observed scalar spectrum is slightly red, ns<1. No choice of a positive decaying α fixes that.
Rigor theorem 5 — local finite-range infrared no-go
For an isotropic equal-time connected correlation C(r) with finite radial moments,
and the Taylor expansion sin(kr)/(kr)=1-(kr)^2/6+O(k^4) gives
Therefore the dimensionless scalar power satisfies
so the infrared limit corresponds to ns→4, not ns≈1. This is a theorem for the stated integrability/moment conditions, not a numerical fit.
Consequence
Rigor theorem 6 — local power-law tails also fail
For 0<α<3 in three spatial dimensions, the distributional Fourier transform obeys
Hence a local tail C(r)∝r−α produces
Every positive decaying local exponent is blue. The red observed branch therefore requires an IR/global object that is not an ordinary local decaying correlation tail.
This failure is scientifically valuable because it eliminates a broad class of stories. But it is still a failure of the claim that ordinary local post-separation influence could replace the large-scale correlation mechanism.
The hidden-global-mode semifinal
Was there a topological or finite global degree of freedom hiding in the frozen Shape that could rescue the spectrum?
A rescue mechanism must correspond to an admitted degree of freedom or a derived effective object. Finite topological labels cannot be rhetorically promoted into a continuum propagating spectrum.
The blind global-kernel audit examined that possibility. The active topology and Actor census do not supply an independent propagating global scalar whose covariance spans a continuum of nonzero Fourier modes. Higher cohomology can classify discrete sectors and finite global constraints can modify a covariance by finite rank, but a homogeneous primordial spectrum is an infinite-rank multiplication operator across a continuum of k.
Translation-invariance theorem
A statistically homogeneous covariance/kernel is diagonal in momentum:
On this is a multiplication operator. If
is nonzero on any set of positive measure, its range is infinite-dimensional. Therefore a nonzero homogeneous continuum kernel cannot be generated by a finite-rank correction.
Theorem 1
A finite number of global constraints or global auxiliaries cannot generate a nontrivial statistically homogeneous power spectrum across a finite band of nonzeroFinite topology inherits the same obstruction
The active compact space has finite Betti numbers. Harmonic representatives therefore form finite-dimensional spaces. Even if each admitted cohomology class were paired with a lawful global coefficient, the resulting sector would contain only finitely many independent global amplitudes.
By Theorem 1, such a sector cannot be the complete homogeneous primordial continuum kernel.
The key linear-algebra point is simple. Conditioning a covariance C0 on N global constraints produces a correction of rank at most N. Changing P(k) over an interval requires infinitely many independent Fourier directions unless the baseline state already carried the required law.
The state-selection comeback
Can the missing long-range covariance arise as a lawful quantum state of existing Actors rather than a new Actor?
The cosmological state must be selected by an upstream admissibility/variational rule. Choosing a covariance because it matches the CMB is fitting, not explanation.
This was the reason to create the State Selection and Boundary Admissibility building block. The block is methodological infrastructure, not a cosmological answer. It prevents four common cheats: “allowed” becoming “selected,” symmetry becoming uniqueness without proof, Hadamard regularity becoming a unique vacuum, and a nonlocal boundary kernel being inserted without parent provenance.
Core theorems and lemmas
Theorem 1 — Permission does not imply selection
If the admissibility conditions define a set S_adm containing more than one operationally inequivalent element, then admissibility alone does not select a physical state. A terminal prediction requires either a further parent-owned selection law or an explicit finite calibration lane.
Theorem 2 — Symmetry does not generally imply uniqueness
For homogeneous/isotropic Gaussian states, invariance usually reduces the covariance to functions of |k|; it does not fix those functions. Therefore the invariant-state set is typically infinite-dimensional until additional dynamics/boundary conditions are supplied.
Theorem 3 — Hadamard is a UV admissibility class, not a unique vacuum
Two Hadamard two-point functions may differ by a smooth function. That smooth freedom can encode distinct infrared occupation and correlations. Hadamard regularity is essential for renormalized local observables, but it cannot by itself determine the primordial infrared kernel.
Theorem 4 — Local finite-derivative boundary kernels have discrete infrared powers
For a rotationally invariant local quadratic boundary action with finite derivative polynomial F(-∇²), the momentum kernel F(k²) is analytic in k². If its first nonzero term is c_m k^{2m}, then the direct Gaussian covariance scales as k^{-2m}, giving dimensionless power k^{3-2m}. A mildly red nearly scale-invariant continuum exponent is therefore not produced by simply choosing a local finite-derivative boundary term.
Theorem 5 — Parent locality does not forbid induced boundary nonlocality
A local bulk operator can induce a pseudodifferential boundary operator after solving the bulk boundary-value problem. Dirichlet-to-Neumann maps and exact Schur complements are therefore lawful sources of nonanalytic momentum dependence. The admissibility question is provenance, not superficial locality of the final boundary kernel.
Theorem 6 — Unique-derived terminal requires zero residual state dimension
After physical and operational quotienting, a unique-derived terminal requires exactly one surviving operational branch and no continuous or functional freedom. If a Bogoliubov function, spectral density, self-adjoint extension parameter, or background branch remains physically distinguishable, uniqueness is OPEN.
Theorem 7 — Calibration can preserve falsifiability
A finite p-dimensional parameter family can be calibrated openly if each independent fitted direction consumes declared calibration information and the vector is frozen before held-out tests. Functional freedom is not finite-dimensional calibration unless a basis/truncation/error theorem proves it.
The companion Boundary-State Test then attacks the candidate routes. A local finite-derivative boundary action has an analytic kernel in k² near the origin. If Ω(k)∝k2j is its leading term, P∝k−2j and the dimensionless tilt is discrete, ns=4−2j. That gives …4,2,0,−2… rather than a generic mild tilt near one. Hadamard regularity fixes UV singularity structure but leaves smooth infrared freedom. Finite-time instantaneous vacua are representation-dependent. A decelerating Big-Bang control background does not give each fixed cosmological mode an asymptotic past WKB region.
The local finite-derivative boundary theorem
Take a translation- and rotation-invariant quadratic boundary action on Σ with a finite number of spatial derivatives and coefficients frozen independently of cosmological targets:
Fourier transformation gives K(k)=F(k²), analytic at k=0. Suppose the first nonzero coefficient is cm, so K(k)~cmk2m.
For a direct Gaussian curvature state with covariance proportional to K⁻¹,
Thus the dimensionless spectral exponent is restricted to the discrete sequence 3,1,−1,−3,… as m=0,1,2,3,… . A continuously adjustable mild departure from scale invariance cannot be generated by the leading term of a finite local derivative expansion.
Decelerating Big-Bang backgrounds fail the asymptotic-past WKB test
Consider a control background a(t)∝tp with 0<p<1, representing decelerating expansion. Then
As t→0⁺, k/(aH)→0 for every fixed comoving k. Thus fixed cosmological modes approach the earliest boundary on the long-wavelength side, not the asymptotic Minkowski/WKB side. A vacuum definition based on k/(aH)→∞ therefore cannot be justified merely by moving the initial slice earlier toward the Big Bang.
This is a statement about decelerating power-law Big-Bang histories. It does not rule out a native background with an asymptotic past, a contracting branch, an accelerated phase, a Euclidean cap, or another parent-owned history that changes the causal mode geometry.
One route survives without target-loading: a local parent bulk history can induce a nonlocal boundary operator through its on-shell Dirichlet-to-Neumann map. The half-space control explicitly demonstrates the mechanism: a local bulk Laplacian generates an |k| boundary kernel. In cosmology the analogous object must be derived from the actual native saddle, not guessed.
Explicit control: a local bulk operator induces |k| on the boundary
Consider a Euclidean massless scalar on the half-space y≥0 with local action
For each Fourier mode the regular bulk solution is
Substitution into the on-shell action gives
The boundary operator |k| is nonanalytic in k² even though the bulk action is strictly local. This control proves the conceptual point: boundary nonlocality is not automatically illicit. What matters is whether the kernel is a derived Dirichlet-to-Neumann map of an owned parent problem.
General Dirichlet-to-Neumann construction
Let L[Φ̄] be the physical quadratic bulk operator around a solved native background Φ̄. Given boundary data Q on Σ, solve
with the parent-selected condition on the complementary boundary/asymptotic support. The canonical normal derivative/momentum defines the Dirichlet-to-Neumann map
The quadratic on-shell action is then
including measure, counterterm, gauge/BFV, and edge contributions required by the Boundary authority. This Hessian is the natural candidate source of the Gaussian boundary-state kernel. It is not chosen mode by mode.
The no-fudge scalar-spectrum test
If we use the simplest surviving scale-invariant state idea, does it match the sky?
The amplitude, tilt, running, tensor content, non-Gaussianity, and isocurvature structure must be generated before those observables are opened, except for any explicitly paid calibration observables.
No. Exact scale invariance predicts ns=1. Using the Planck scalar-tilt comparator ns=0.9649±0.0042, the discrepancy is
That is too large to describe as “basically right.” The point of the test is not that Planck must remain the final word forever; it is that a prospective exact-scale-invariance branch already has a clear held-out comparator and loses badly.
The pump-field form of a standard canonical scalar perturbation equation illustrates the required structure. If z″/z=(ν²−1/4)/η² and the state is selected in the appropriate asymptotic branch, then ns−1=3−2ν. The measured comparator corresponds to ν≈1.51755 and ν²−1/4≈2.052958. Those are inverse targets after unblinding, not predictions.
Post-unblinding inverse targets—not predictions
From the pump-field control relation ns=4−2ν, the measured central value corresponds to
If curvature itself were represented directly by a Gaussian boundary kernel with no additional scale-dependent transfer, then
These equations are diagnostic inverse requirements. A future Shape calculation is not allowed to choose ν=1.51755 or exponent 3.0351 because these values look correct. It must derive its own background and kernel and then be compared.
The vacuum-stability red-zone match
Does the selected internal geometry actually sit at a stable physical vacuum?
The selected physical background must be a lawful quantum vacuum or metastable state on the relevant timescale. A forced physical tachyon or uncontrolled negative Hessian direction is branch-level evidence against the frozen solution.
This is the most important non-cosmological falsifier in the current authority. The location u=(1,1,1) is symmetry-forced as a critical point, but criticality is not stability. The source computes negative shape-doublet curvature contributions and identifies a physical fixed-volume shape mass-squared of −1/3 in the relevant geometric reduction. The opening owner-item labels the branch LIVE-FALSIFIER-AS-WRITTEN.
The detailed SG-6 dossier is nuanced: it keeps tree-level, geometric-curvature, bosonic-zeta, and fermionic-supertrace layers distinct, and it says the full physical Hessian depends on an unpinned fermionic contribution. That nuance prevents us from asserting that the full all-order vacuum is definitely unstable. It does not allow us to advertise a stable vacuum.
GATES_SOURCE_OF_TRUTH.md, lines 6–8; 12063–12083; 13132–13148The selected center is a critical point. Two computable contributions point toward a saddle. The remaining fermionic sign/weight is not fixed by the frozen shape. The opening owner-item explicitly calls vacuum stability a live falsifier as written.
If the compactification does not define a metastable physical background for long enough to support the low-energy spectrum, then downstream flavor and cosmology calculations are being performed around the wrong background. The correct order is therefore to settle the net physical Hessian before treating the active branch as a candidate universe.
Stop condition. The branch clears this round only if the full, regulator-consistent physical Hessian—including the fermionic graded term and the declared renormalization prescription—is derived prospectively and is positive on every physical modulus direction needed by downstream calculations. If it remains negative, the frozen branch loses.
The flavor prediction upset
Do the forced flavor outputs agree with data closely enough to support the frozen one-angle ansatz?
A claimed forced flavor relation must survive same-ruler comparison with precision data. A confirmed ~14σ forced discrepancy cannot be converted into a win by status taxonomy.
The flavor gate contains genuine successes, but one standing prediction is too discrepant to average away. The forced one-angle CKM construction gives |Vtd|=0.01145 against 0.00857±0.00021, a pull of
Matching the central value would require multiplying that entry by approximately 0.7485, a roughly 25.2% downward correction. The same gate also carries a forced CKM phase of +60.0° against 65.5°±1.5°, a 3.67σ tension.
GATES_SOURCE_OF_TRUTH.md, lines 17490–17498; 18940–18946; 19476–19486The source explicitly retracts an earlier theory-error band that had masked the tension and now reports the ~14σ |Vtd| residual and 3.67σ CKM-phase residual against experiment.
This matters epistemically. A 14σ discrepancy is not made small because five other CKM magnitudes are close. If all eight non-anchor entries were independent under the ansatz, the outlier would be overwhelming evidence of model misspecification. They are not fully independent, but that makes the conclusion stronger in one sense: the one-angle structure couples them, so a repair must preserve the successful entries while moving the failed one.
A subleading texture/operator correction is scientifically allowed only if its operator form and coefficient provenance are derived from the frozen geometry or from a new building block before the repaired |Vtd| value is evaluated. Because the target is already known, the cleanest test is to derive the correction from other sectors/constraints and then evaluate a held-out flavor vector, not tune directly to |Vtd|.
The baryogenesis endurance round
Does the theory derive why the universe contains more matter than antimatter?
If the theory claims cosmological completeness, the observed matter-antimatter asymmetry must arise from admitted dynamics or be openly charged as an external input. Naming baryogenesis as a certified wall is not a cosmological prediction.
The current gate authority does not claim that it does. BG-10 reaches a certified-irreducible terminal because the seesaw sector has an exact rescaling flat direction:
that leaves the low-energy neutrino sector invariant while moving the absolute heavy-Majorana scale required by the leptogenesis calculation. The gate also records a separate C-odd sign/orientation issue. In the project’s governance taxonomy this is a legitimate terminal. In the theory tournament it is simply not a prediction of ηB.
GATES_SOURCE_OF_TRUTH.md, lines 34923–34977; 35230–35382The dossier says the absolute heavy Majorana scale cannot be recovered from the current fixed record because of an exact flat direction, and describes the baryogenesis result as CERTIFIED-IRREDUCIBLE rather than numerically derived.
This is not evidence that the rest of the framework is false. It is evidence against calling the current framework a complete cosmology. State Selection and Boundary Admissibility may eventually change the ownership question if it derives a boundary asymmetry or heavy-sector normalization from a genuine parent rule, but that would be a new result and must not be back-credited to BG-10.
The dark-sector possession test
Does the frozen branch supply the gravitational source normally attributed to dark matter?
The framework must supply the stress-energy and perturbation behavior that explain the dark gravitational phenomenology, or explicitly narrow its claim. Matching only a present density is insufficient.
The Gap-11 authority finds a serious candidate route but explicitly states that the relic normalization depends jointly on geometry-side portal normalization and cosmological thermal history/reheating. One observed relic abundance therefore constrains at least two unknown structures. The gate classifies the abundance as anchor-limited by TRH.
GATES_SOURCE_OF_TRUTH.md, lines 37721–37857; 37791–37797The canonical endpoint is CLOSED / CERTIFIED-IRREDUCIBLE + finite portal-compute forecast. The relic abundance is not a pure output of timeless geometry; it is a mixed geometry + cosmological-history record.
For cosmology, that means the native background still lacks a derived dark clustering stress tensor and abundance history. A candidate particle identity is not enough: the theory must supply ρX(a), pressure, perturbation sound speed/anisotropic stress where relevant, production history, and the transfer consequences for CMB peaks and structure growth.
The home-field cosmology test
Can the theory derive its own background expansion rather than borrowing ΛCDM?
The cosmological background and perturbation transfer must be derived from the parent/effective action rather than imported from a best-fit ΛCDM history. The same background must support BBN, CMB, BAO, growth, and age.
The final cosmology cannot be declared native until the compactified parent action is evaluated on a cosmological ansatz, all retained homogeneous fields/moduli are included, constraints are varied consistently, and the resulting solution determines H(a). Importing H0, Ωm, ΩΛ, or a desired age would turn the age and distance predictions into reconstructions.
Rigor theorem 10 — native background reduction must be variational
A cosmological background cannot be declared by analogy. Starting from the complete parent action, insert a homogeneous/isotropic retained metric while preserving every frozen internal equation or reaction field:
The background equations are then the complete Euler–Lagrange vector, including reaction equations associated with frozen internal variables:
Until that vector has a solved branch, HShape(a) remains OPEN. Substituting the standard Friedmann equation is allowed only in the segregated control lane.
The actual missing parent calculation
To execute the surviving route, the theory must first supply a native homogeneous/isotropic (or explicitly alternative) saddle Φ̄. The correct object is obtained by reducing and varying the parent effective action, not by importing H(a) from ΛCDM:
Then one must compute the complete Hessian, constraints, gauge reduction, and physical scalar block:
where the second equation is the schematic Schur complement after nondynamical/constraint variables C are eliminated. Interdependence explicitly requires all off-diagonal sectors to be computed, proved zero, or integrated out with controlled Schur complements.
The state-selection problem is coupled to this round. The physical scalar perturbation action is background-dependent, and the lawful nonlocal boundary kernel must be the on-shell Hessian/Dirichlet-to-Neumann map of that solved background. Therefore the “native background” and “primordial state” are not two independent missing boxes; they are one joint calculation.
The referee audit
Does Constraint-Based Reconstruction itself help us find truth, or can its governance language protect a favored theory from losing?
A reviewer who cannot repair the theory must be able to reproduce the decisive calculations, find known defects, and obtain the same pass/fail result from frozen artifacts.
The approach has several unusually strong features: explicit assumption sweeps, freeze-before-compare, same-ruler audits, candidate-class exhaustion, destructive controls, negative terminals, provenance hashes, and willingness to publish closed-negative results. Those features are exactly why the local-cosmology and finite-global-mode routes were eliminated instead of quietly reparameterized.
But the tournament exposes one methodological danger: the project’s terminal taxonomy is optimized to ensure every research question ends in a well-typed state. That is excellent project management, but a phrase such as “33 RESOLVED +0 / 0 OPEN” can sound like “33 physical problems solved.” It does not mean that. CERTIFIED-IRREDUCIBLE may mean “the theory cannot determine this quantity”; CLOSED-NEGATIVE may mean “the proposed mechanism fails”; a measured-anchor terminal may mean “the value is consumed from experiment.”
Maintain two permanent public ledgers:
- Closure ledger: whether each gate has reached a legitimate research terminal.
- Truth/viability ledger: empirical passes, empirical tensions, closed-negative mechanisms, live falsifiers, calibrated quantities, and unresolved physical predictions.
No aggregate “all gates closed” statement should appear without the second ledger beside it.
With that correction, the method earns a strong mark. A methodology that can expose a 14σ miss, retire an incorrect uncertainty band, prove its own local mechanism cannot work cosmologically, and identify a live vacuum falsifier is doing real epistemic work.
The championship conditions
What would actually make us say “the theory is probably right”?
The evidence must be jointly overconstraining: one consistent branch and one consistent parameter ledger must survive all decisive domains simultaneously. Separate local wins do not sum to a unified-theory win if they require incompatible branches.
A complete theory of nature is never proved by a finite collection of empirical tests in the mathematical sense. What we can demand is a sequence of increasingly expensive, prospectively frozen predictions that would be very unlikely to succeed by flexible fitting.
| Championship test | Pass condition | Immediate loss condition |
|---|---|---|
| Vacuum stability | Full physical modulus Hessian is positive/metastable at the selected branch under a frozen regulator/renormalization prescription. | A physical tachyonic direction remains with no independently derived stabilizer. |
| Flavor replay | A prospectively derived correction/operator reproduces held-out flavor observables while preserving existing successful entries and using no Vtd-targeted coefficient. | Repair coefficient/form is selected from the known Vtd miss or spoils other forced predictions. |
| Native background | Parent action produces a self-consistent H(a), conserved stress ledger, and acceptable phase history without importing cosmological best fits. | No physical saddle exists or required history is inserted as a free function. |
| Boundary state / primordial spectrum | Parent-induced DtN Hessian fixes the state kernel; scalar amplitude/tilt/running emerge before CMB comparison. | Kernel or tilt is inserted from CMB, or the native state produces a strongly excluded spectrum. |
| Tensor/non-Gaussian/isocurvature | Same state/background predicts acceptable r, fNL, and isocurvature without new dedicated tuning. | Additional arbitrary knobs required for each statistic. |
| Thermal history | Derived background + particle content survives BBN/recombination and predicts held-out abundances/spectra. | BBN or CMB failure outside declared systematic uncertainty. |
| Late universe | Same frozen lane predicts BAO/growth/age with a small paid calibration budget. | Independent free function required to follow each dataset. |
The championship criterion is deliberately multiplicative rather than additive: a spectacular win in one round does not cancel a fatal loss in another. A viable universe has to survive all load-bearing sectors simultaneously.
Would I call the theory correct today?
No. Would I stop the research program? Also no.
The final public claim must be exactly proportional to the strongest level actually cleared: method merit, viable candidate, predictive theory, strong evidence, or best-supported theory within scope. No stronger wording is allowed.
The most defensible verdict after this tournament is asymmetric:
| Object | Verdict today | Reason |
|---|---|---|
| Constraint-Based Reconstruction as a research method | MERITORIOUS / CONTINUE | It enforces useful anti-fitting structure and has already generated genuine negative results and corrections. |
| Microscopic finite-record / first-separability mechanism | PASS-CONDITIONAL | Internally coherent at the stated finite-region/quadratic-envelope scope; absolute time inherits a pre-existing particle-physics calibration. |
| Frozen 13D Shape as the correct physical branch | NOT CLEARED | Vacuum stability is a live falsifier as written, and the one-angle flavor ansatz carries a ~14σ forced discrepancy. |
| Local-interdependence explanation of primordial structure | REJECTED | Finite-range and ordinary decaying local correlations predict the wrong infrared spectrum. |
| Finite hidden global/topological rescue | REJECTED | Finite-rank global structure cannot generate the required continuum covariance. |
| State-selection / parent-induced DtN route | LIVE BUT OPEN | Mathematically legitimate route; native background and state kernel not yet derived. |
| Baryogenesis | NOT PREDICTED | Exact flat direction/scale ownership wall in current branch. |
| Dark sector | NOT PREDICTED NATIVELY | Candidate portal exists, but relic abundance and thermal history are anchor-limited. |
| Full cosmology | NOT CLOSED | No native background + state solution yet; downstream agreement controls are not native predictions. |
The approach has earned the right to continue. The current frozen theory has not earned the right to be called correct.
That is not a disappointing terminal. It is exactly the distinction a serious research program needs. The method is useful if it tells us when a beautiful branch is wrong. The next phase should therefore be narrow and ruthless: stop broad cosmology expansion until the two red-zone branch tests are addressed prospectively.
The reviewer’s current scorecard
On the evidence presently recorded, the research method is ahead of the physical theory. That is not an insult; it is what a falsification-capable method should sometimes report.
| Burden | Current status | What remains before promotion |
|---|---|---|
| Research method catches errors | PASS | Continue independent hostile review; do not weaken after successes. |
| Mathematical viability of selected branch | RED ZONE | Resolve the physical Shape-doublet vacuum Hessian/stability question. |
| Complete parameter/provenance ledger | PARTIAL | Threshold normalization provenance and all injected reals/branch choices must remain charged consistently. |
| No forced precision contradiction | RED ZONE | Adjudicate the ~14σ |Vtd| result: optional ansatz failure or forced-branch falsifier. |
| Microscopic record theorem | CONDITIONAL PASS | Maintain scope: first-existence finite-record theorem, not universal factorization. |
| Primordial local mechanism | CLOSED-NEGATIVE | Do not reuse local tails as an inflation replacement. |
| State selection / primordial kernel | OPEN-FUNCTIONAL | Derive native saddle and parent-induced Dirichlet-to-Neumann state kernel without CMB tuning. |
| Baryogenesis | NOT NATIVE PREDICTION | Derive asymmetry or narrow the completeness claim. |
| Dark sector | NOT NATIVE PREDICTION | Derive stress-energy and perturbation behavior or narrow the claim. |
| Native cosmological history | OPEN | Derive H(a), perturbation transfer, thermal history, and late-time observables from one branch. |
| Prospective distinctive prediction | NOT YET ADEQUATE | Freeze a prediction that was not used in branch construction. |
| Independent reproduction | PARTIAL | External replay of decisive chains from frozen artifacts. |
The order of operations: where the goalposts say we should spend effort
Once the finish line is fixed, the priority order changes. We should not work on the most exciting calculation; we should work on the calculation with the highest decision value.
- Vacuum stability first. Compute the full physical Hessian—including admissible loop/Casimir completion, gauge-zero-mode removal, scheme dependence, and uncertainty—to decide whether the selected Shape background actually exists as a viable state. A branch that cannot sit where the rest of its calculations assume it sits should not be sent deeper into cosmology.
- Flavor falsifier second. Determine whether |Vtd| is genuinely forced by immutable structure or by an auxiliary one-angle ansatz. If forced and the same-ruler discrepancy survives independent replay, reject the branch. If auxiliary, retire the ansatz and prospectively select the next lawful flavor mechanism before re-opening the target.
- Only then run native background + state selection. Solve the 4D effective cosmological saddle, derive the physical perturbation operator, and obtain the boundary kernel from the on-shell/Dirichlet-to-Neumann map. This is the correct test of whether the cosmology emerges rather than being inserted.
- Then fill the cosmological matter ledger. Baryogenesis and dark-sector stress-energy must be derived or explicitly calibrated before a full H(a) claim is scored.
- Finally run the blinded end-to-end cosmology. BBN, CMB, BAO, growth, age, and any novel feature are opened only after the upstream packet is frozen.
It prevents a seductive cosmological match from making us psychologically reluctant to accept a fatal particle/vacuum result. The cheapest decisive falsifier gets the ball first.
What would actually convince me?
If the following sequence happened without moving these goalposts, my assessment would change materially.
- The full physical vacuum calculation removes the current live stability falsifier from frozen ingredients, or shows a controlled metastable lifetime far exceeding the required physical history.
- The flavor discrepancy is either prospectively repaired by a lawful mechanism selected without |Vtd| target access, or demonstrated not to be a forced output of the core branch. The resulting full CKM/PMNS package survives a precision replay.
- The native cosmological saddle exists and the State Selection and Boundary Admissibility rules select a unique or finitely classified physical state without CMB input.
- The induced primordial spectrum has the correct qualitative class—near scale invariance, mild red tilt, acceptable tensors/non-Gaussianity/isocurvature—before those data are opened, with any calibration observables explicitly removed from the prediction score.
- The same branch yields a viable baryon ledger, dark-sector gravitational behavior, and thermal history, then predicts BBN, CMB peak structure, BAO/growth, and age without mutually inconsistent re-tuning.
- The joint held-out predictive score is competitive with or better than reasonable benchmark models after paying the theory’s effective complexity.
- At least one genuinely distinguishing prediction is locked before a new or genuinely unused datum and succeeds.
- An independent reviewer replays the decisive chain and reaches the same conclusion.
If all eight happened, I would no longer describe the project merely as an interesting reconstruction program. I would describe it as strong evidence that the framework is capturing real physical structure. That is the correct finish line. Anything less should be labeled by the lower rung it actually clears.
What would not convince the reviewer—even if it looks impressive
- “All gates are resolved.” Governance completeness is not empirical truth.
- “The model reproduces many known numbers.” Not unless the parameter/branch search cost and target visibility are accounted for.
- “The required mechanism exists mathematically.” Existence is weaker than selection.
- “A calibrated model matches the calibration data.” That is what calibration means.
- “The theory has fewer named parameters.” Not if a free function, branch grammar, or hidden normalization carries equivalent flexibility.
- “The discrepancy is a certified falsifier and therefore a strength.” Honest reporting is a methodological strength; the contradicted physical mechanism still loses.
- “The cosmology works if we assume the primordial kernel/background.” That may validate a downstream transfer engine, but it does not establish a native cosmology.
- “No one has disproved it.” Absence of disproof is not positive evidence.
- “The framework is elegant or unifies language.” Elegance is a prior preference, not a likelihood.
The NASA-facing championship evidence packet
If we ever send the strongest version of this work to NASA or another technically demanding external group, the package should be designed so that they can audit it without contacting us.
The reader should receive one folder that answers four questions: What exactly is the theory version? What did you fit? What did you predict before seeing the answer? Can I run it myself?
The championship packet should be content-addressed and immutable. It contains: theory/geometry manifests; input provenance; parameter/function/branch ledger; calibration contract; prediction hashes/timestamps; complete source code and environment lockfile; machine-readable input/configuration files; raw/processed data references; exact same-ruler observable maps; uncertainty and nuisance specification; negative controls and ablations; benchmark model implementations; independent replay output; and a one-page claim graph linking every headline statement to its executable witness.
Required packet contents
- One-page simple brief. Premise, claim, what was fitted, what was predicted, what would falsify it.
- Technical claim ledger. One row per claim with exact equation, scope, status, provenance, and witness hash.
- Frozen theory manifest. Geometry, Actors, Rulebook, boundary/state rules, and version hashes.
- Complete flexibility ledger. Scalars, functions, branches, priors, search grammar, rejected candidates, and calibration observables.
- Reproducible computation. Source, dependencies, seed policy, configuration, numerical tolerances, and deterministic/repeated-run expectations.
- Blind prediction packet. Timestamped outputs and scoring rule created before target access.
- Benchmark packet. Reasonable competing models run through the same observation and nuisance pipeline.
- Independent replay. Separate implementation or team, including discrepancies and adjudication.
- Failure registry. Every known red zone, closed-negative branch, and repair history.
The NASA open-science references above motivate transparency and reproducibility; this championship packet is intentionally stricter than those general process requirements because the scientific claim is unusually ambitious.
The prediction packet we should freeze before the next major test
Every decisive future calculation should produce a small machine-readable packet before target comparison:
{
"theory_manifest_hash": "...",
"code_hash": "...",
"branch_id": "...",
"measured_anchors": [...],
"fitted_parameters": [{"name":"...","calibration_datum":"..."}],
"discrete_choices": [...],
"external_theorems_or_floors": [...],
"prediction_observables": [...],
"prediction_values_or_distributions": [...],
"same_ruler_map": {...},
"uncertainty_model": {...},
"negative_controls": [...],
"fatal_falsifiers": [...],
"timestamp_or_commit": "..."
}The packet is more important than another hundred pages of prose because it makes later historical revision difficult. The prose explains; the packet proves what was frozen.
Independent reviewer protocol
The project’s dossier-review agent already contains a valuable discipline: a reviewer finds defects but does not repair them. The theory-level review should extend that separation.
- Give the reviewer the frozen theory manifest, decisive dossiers, evidence artifacts, code, and prediction packet.
- Require arithmetic/algebra replay, internal consistency, cross-document consistency, evidence provenance, falsifier quality, registry completeness, and claim-evidence proportionality.
- Add theory-level checks absent from a single-gate review: parameter-search accounting, branch-fatal classification, benchmark comparison, held-out predictive score, and cross-domain mutual compatibility.
- Run known-answer tests: the reviewer must catch the vacuum-stability red flag, the flavor tension, and the distinction between governance closure and truth. A reviewer that misses known defects is not trusted on unknown defects.
- Publish the defect report before repair. Then repair in a separate process and rerun the same reviewer.
Independent review does not mean the reviewer must agree with every modeling choice. It means the factual chain from declared choices to claimed consequences is reproducible and the claim strength is not inflated.
The no-moving-goalposts decision tree
Once this tree is published, a future result changes the score, not the tree. That is the point.
The only claim worth fighting for
If we clear this edition’s goalposts, the strongest defensible public statement is not “we proved everyone else wrong.” It is stronger scientifically:
That is the kind of result a serious scientist cannot responsibly ignore. They may still disagree, but they must engage the evidence.
Even this statement remains conditional and falsifiable. Future measurements can reject the branch. “Strong evidence” is the endpoint of this tournament; logical certainty is not available for an empirical theory.
Final whistle: where the goalposts really belong
The correct goalpost is not “derive everything,” and it is not “close every gate.” It is survive enough independent opportunities to be wrong that the remaining success is difficult to explain by flexibility, calibration, coincidence, or bookkeeping.
For this project, that means first clearing the two existing branch-level red zones; then deriving—not choosing—the native cosmological background and state; then producing a paid, overconstrained, cross-domain prediction vector; then surviving independent replay and at least one genuinely distinguishing prospective test.
That is a demanding standard. It is also achievable in principle. Most importantly, it tells us in advance what result would make us stop, what result would merely narrow a claim, and what result would actually deserve increased belief.
Appendix A — Frozen burden-of-proof matrix
| Claim | Minimum evidence | Disqualifier |
|---|---|---|
| “Method has merit” | Reproducible calculations, negative controls, documented failures, defect-catching review. | Review process systematically launders failures or cannot reproduce its own arithmetic. |
| “Branch is viable” | Stable lawful state/background, consistency, no confirmed forced contradiction. | Physical instability, anomaly/ghost, or core-forced high-significance empirical contradiction. |
| “Theory predicts X” | X excluded from calibration and branch selection; frozen derivation/provenance. | X or close surrogate used to choose parameter/function/branch. |
| “Theory unifies domains” | One branch/ledger simultaneously predicts independent domains. | Different incompatible branches or domain-specific arbitrary functions. |
| “Strong evidence theory is right” | Above + novel distinguishing prediction + independent replay + benchmark competitiveness. | Success disappears out of sample or under independent implementation. |
Appendix B — Parameter and flexibility constitution
Every theory element that can move an observable belongs to exactly one class:
- Measured root/anchor: imported measurement; never scored as predicted.
- Frozen structural choice: discrete geometry/category/branch choice fixed before the tested observable; selection history must be disclosed.
- Derived quantity: deterministic consequence of earlier classes.
- Fitted scalar parameter: charged against calibration data; predictions begin after freeze.
- Free function/kernel: charged by its basis/effective dimension or prohibited from “parameter economy” claims.
- Nuisance parameter: measurement/instrument description; marginalized but not confused with fundamental theory freedom.
- External theorem/floor: mathematical/physical fact the framework consumes rather than derives.
No object may change class after target comparison without an explicit versioned amendment. This is especially important for threshold vectors, normalizations, vacuum terms, and state kernels.
Appendix C — Source integrity ledger
The following local artifacts were used to construct this review. Hashing them does not prove their claims; it freezes which texts were consulted.
| Artifact | SHA-256 |
|---|---|
GATES_SOURCE_OF_TRUTH.md | f85cdae7e4918681b29e35568e7a3d1f3841d739a3cad2c48eb28a0dc5888b90 |
GUT.md | f1fb418c93f004afac1d307ee585a5b1390949d84e3496df77353ecaba89e658 |
HIKING_PHYSICS_MASTER_IMPLICIT_ASSUMPTIONS_LEDGER_v1_3(2).md | 77ddd3217d6aecfd3c9dc03cbe8ac8dd1ab73eaf45adf06908dd181579f415f9 |
handoff-dossier-review-agent.md | 80f2bf731f50e41388754e9b1eaff6b60a074cd762c7905df996a3898eb231ba |
State_Selection_and_Boundary_Admissibility_Max_Rigor.html | 39813b59ad957be8796b8f6ddfe82e9d59921a61250c5b067a8f0eca83a1dcc8 |
Boundary_State_Primordial_Spectrum_Max_Rigor_Test.html | ebefb119096e690f343efb8124fc9272a7f4d907bc33ad29bef2b3c52a50f0a9 |
Constraint_Based_Reconstruction_Decisive_Theory_Tournament.html | 7d02b2501b178e16f1101606e98e0aa02d25b7567f3c2101b3977d1c1b8859a9 |
External reproducibility standards consulted for this edition
- NASA Open Science — transparent, available, reproducible, collaborative science.
- NASA Science Information Policy (SPD-41a) — public availability of publications/data/software and reproducibility-oriented guidance.
- NASA SPD-41a FAQ — software and input/configuration files needed to reproduce/validate results should be archived appropriately.
- NASA Guidelines for Promoting Scientific and Research Integrity — peer review, disclosure of assumptions/biases, open sharing, honest reporting.
These sources set process expectations only; they are not cited as scientific evidence for Constraint-Based Reconstruction.
Appendix D — Reviewer checklist before any future “we are done” claim
- What exact sentence are we trying to promote?
- What is the strongest lower-level claim already established?
- What new burden distinguishes the desired claim from that lower level?
- Which data were visible when the branch, parameterization, and candidate grammar were chosen?
- Which numbers are measured anchors, which are fitted, and which are genuine predictions?
- Are all free functions charged by effective flexibility?
- Does the selected physical background exist and remain stable?
- Are there any confirmed forced contradictions?
- Is the prediction vector jointly overconstraining rather than a collection of correlated plots?
- Does one branch survive all domains?
- What reasonable benchmark is being compared?
- Do destructive controls kill the result in the predicted manner?
- What outcome would make us reject the branch?
- Was that outcome written before the calculation?
- Can an independent implementation reproduce the decisive step?
- Is there at least one high-value prediction not available during model construction?