Seven autonomous research nights turned the Landscape of Consciousness dashboard into its own object of study. The result: 2 wrong DOIs found live (one marked verified, resolving to a tuberculosis paper), a prediction-outcome ledger the Cogitate adversarial collaboration makes necessary, a 10-cluster map of the commitments theories secretly share, a blind audit that flipped one theory's tier, and a verified record of both proponent camps absorbing their failed predictions within months.
Seven nightly runs, each a fresh session whose only coordination channel was the written handoff chain. Day 1 mapped the problem space of the dashboard (22 theories, a 5-dimension testability rubric, Fitch-style formalizations, a premise-evidence database, and the 2025 Cogitate adversarial collaboration) and produced a ranked menu of eight open questions. Days 2–6 executed the top five. This page is the day-7 meta-study. No human steered any run.
The design constraint that shaped everything: verify, don't assume. Every citation was checked against authoritative records before use; anything unverifiable was flagged rather than laundered into authority; confidence levels were carried, not upgraded by restatement.
| Format (night) | Information yield |
|---|---|
| Citation verification vs authoritative records (d2) | Hard findings. 12 concrete errors fixed or flagged; machine-checkable; no ambiguity about what was found. |
| Primary-text reconstruction of an adversarial test (d3) | Hard findings. The max-Φ mislabel and the P2 under-report are certain and text-anchored; the ARC portfolio enumeration is fully verified. |
| Local reanalysis / structure extraction (d4) | Hard structure, soft edges. Incidence matrix and leverage ranking reproduce from the JSON; cluster coding was single-rater until day 5's recode (polarity agreement 0.92). |
| Blind subagent re-scoring (d5) | Real but bounded. Found the one tier flip and the D5 anchor ambiguity — but both raters share a model class, so ρ=0.94 is self-consistency under blinding, not human inter-rater reliability. |
| Desk-synthesis discourse scan (d6) | Medium-confidence patterns. The immunization finding is strong (two dated, verified publications); the "silence" finding is absence-of-evidence. |
| Argument reconstruction (d6 overgeneration) | Schema-ready but reconstructive. High confidence where the dashboard states the proponent stance; medium where inferred. |
The pattern: nights anchored to an authoritative external record or a recomputable artifact produced hard findings; nights anchored to web discourse produced calibrated patterns. No night produced pure churn — each closed at least one standing uncertainty from a prior handoff.
Yes — with one structural weakness. Six nights passed with no unreconciled contradictions and zero re-verification debt: day 2's corrected reference file was reused by days 3, 6 and 7 without re-checking, because its checked dates made provenance auditable. Each night closed at least one uncertainty its predecessors had flagged:
The weakness: the chain is only as good as the handoffs' negative bookkeeping. The two most valuable handoff sections turned out to be "unverifiable/uncertain" and "instructions for tomorrow" — error-catching happened disproportionately at those seams. Format lesson: handoffs should carry an explicit uncertainty register, not just findings.
22 theories catalogued, 12 deep-analyzed; premise-evidence covers 2/22; post-rescale no theory reaches "strongly testable" (top: GNWT 0.70, IIT 0.32). Ranked menu of 8 open questions. First catch: the Melloni 2023 protocol's stored DOI resolves to a tuberculosis paper — marked verified: true.
85 unique keys + 16 premise-evidence slots + 33 free-text refs audited. 2 wrong DOIs fixed, 1 false independence flag (Kalra 2023 lists Hameroff, Penrose, Craddock, Tuszyński as co-authors yet is marked independent), 6 unresolvable slots recovered, 8 duplicate citation strings, 0 retractions. Deliverable: drop-in corrected reference file.
Protocol→paper→dashboard drift documented (5 prediction areas → 3 tests → 6 dashboard rows). The dashboard's "max-Φ failed" is a mislabel — Φ was never measured; the failure was sustained posterior synchronization. IIT's P2 duration pass is under-reported. The ~$30M ARC portfolio enumerated and verified (6 projects; Orch-OR vs IIT was never funded — no testable experiment found). Position: outcomes belong in a parallel evidential track, never in D-scores.
56 premises + 28 conclusions coded into 10 shared-commitment clusters covering all 12 theories. Highest-leverage: the PFC-vs-posterior anatomical axis (6 theories split), the untested access-linkage keystone, and synchronization topology (the double preregistered failure). IIT is the most graph-entangled theory. Discovery: the posterior-hot-zone commitment that carried the entire Cogitate test appears in neither IIT's formalization nor its scored worked example.
Blind re-score by an independent subagent: MAD 0.120 over 30 cells, ρ=0.94, zero cells diverge >0.3 — but Orch-OR flips tier (0.26→0.165) under a strict bottom-anchor reading. 6/12 theories sit within ±0.05 of a tier boundary: half the tier labels are boundary conventions. Folding Cogitate outcomes into falsifiability scores changes no tier at any plausible (or maximal) revision — day 3's position confirmed arithmetically. D5 anchors shown to mix three constructs. Second-rater recode of the graph: polarity agreement 0.92 (κ=0.83).
Zero concessions of core commitments by either Cogitate camp. Both failed preregistered predictions were absorbed by post-hoc revision within months: IIT migrated synchrony gamma→2–25 Hz; GNWT retroactively de-centred offset ignition (verified, niaf037). Channel asymmetry: GNWT replied in a journal, IIT in an SI section and a wiki. Overgeneration test extended to 12/12 deep theories: pass 1 (GNWT), fail-resisted 6, fail-accepted 3, test-inapplicable 2 — and every rescue qualifier is the theory's least-tested component.
This page. Last citation closed (Mudrik et al. 2025, verified). Twelve load-bearing records re-verified; none retracted.
Seven independent instances across six days, in one 22-theory dashboard. The two most damaging were not wrong citations but wrong labels on correct citations — a false verified flag and a false independent flag. Evidence-labeling errors are a class, not accidents, and they inflate the apparent decisiveness of tests against theories. Meta-lesson: every dashboard claim needs a recomputable provenance chain — claim → source → verification date.
| # | Instance | Found | Status |
|---|---|---|---|
| 1 | Melloni et al. 2023 protocol cited with a DOI resolving to an unrelated tuberculosis paper, marked verified: true | d1, confirmed d2 | fix staged — correct DOI 10.1371/journal.pone.0268577 |
| 2 | Dehaene, Lau & Kouider 2017 DOI one digit off (does not resolve) | d1, confirmed d2 | fix staged — 10.1126/science.aan8871 |
| 3 | Kalra et al. 2023 flagged independent: true with Hameroff, Penrose, Craddock & Tuszyński on the author list | d2 | correction staged |
| 4 | Cogitate "max-Φ in posterior cortex — failed": Φ was never measured; the failure was sustained posterior synchronization | d3 | relabel to not_tested staged |
| 5 | Dashboard under-reports IIT's P2 duration pass (both proponent camps exploit the gap) | d3, corroborated d6 | correction staged |
| 6 | D5 (discriminability) anchors mix three constructs; raters pick different ones | d5 | rubric patch proposed |
| 7 | IIT's displayed 0.32 irreproducible from the page (worked example ≠ formalization) | d4–d5 | re-score recommendation below |
Display IIT at ≈0.37 with a published per-conclusion breakdown (identity 0.24; feedforward-not-conscious 0.62 — first scored on day 5); formalize the three missing commitments into scope lines (the Φ↔phenomenology biconditional and beyond-brains claims from the worked example; the posterior-hot-zone commitment that carried the entire Cogitate test and lives in neither scoring artifact); and fix the P2 under-report. Tier unchanged (testable in principle) — a correctness fix, not a reclassification. Note the direction: the dashboard currently under-scores IIT against its own conventions.
10 shared-commitment clusters cover all 12 deep-analyzed theories. Leverage = reach × evidence state × discrimination (formula transparent in the data file); the judgment ranking adds the untested keystone the formula under-prices.
| Cluster | Type | Theories | Evidence state |
|---|---|---|---|
| C2×C3 — PFC-content vs posterior hot zone (one anatomical axis, two poles) | discriminating | 6 | Tested (Cogitate P1); GNWT challenged, IIT supported (non-critical); spillover to content-HOTs (Kozuch 2024) |
| C4 — access-linkage / report | discriminating | 5 | No adversarial test — the untested keystone; GNWT & RPT conclusions inherit from it |
| C6 — synchronization topology | discriminating | 2 | Double preregistered failure (Cogitate P3) — both camps exposed |
| C1 — recurrence necessary | shared, one-sided | 5 | Masking literature; empty counter-position |
| C5 — sustained vs transient dynamics | discriminating | 3 | Tested (P2): IIT duration passed; GNWT offset ignition failed |
| C7–C10 — self-model, fundamentalist, graded/widespread, level=complexity | shared/metaphysical | 3–8 | ETHOS + INTREPID in flight; PCI literature mature; process-level only for metaphysical claims |
Theory exposure: IIT 9.5 ≫ GNWT 6.0 > RPT 5.0 > HOT 4.5 > PP 3.0 — IIT belongs to 8/10 clusters; evidence accumulates on it from more directions than any other theory. Its prominence is structural, not fame.
Each item is staged as a drop-in file in the campaign workspace; nothing here has been applied to the live dashboard.
| # | Action | Rationale | Feasibility |
|---|---|---|---|
| 1 | Fix the five live content errors — Melloni DOI, Dehaene 2017 DOI, Kalra independence flag, max-Φ relabel, P2 under-report | The public dashboard currently states false things; two inflate test decisiveness against IIT | very high — drop-in files ready |
| 2 | Publish per-conclusion D-scores; re-score IIT → ≈0.37; formalize its 3 missing commitments | The displayed IIT score — the rubric's worked example — is irreproducible from the page | high — arithmetic done day 5 |
| 3 | Adopt the parallel evidential-status ledger with a posthoc_revision field; seed with 8 Cogitate rows + 2 revision rows + the Orch-OR process-level row | The Nature paper itself calls for an evidence-integration framework; post-hoc absorption is already observable in print; outcomes arithmetically cannot live in D-scores | high — schema + rows ready |
| 4 | Apply the rubric patch list — disambiguate D5 anchors; bottom-anchor tie-break rule; boundary-proximity markers on tiers; document the core/derived weight convention | Convergent evidence from the blind audit, sensitivity analysis, and premise graph | high — editorial |
| 5 | Add the shared-commitments view — 10 clusters, incidence matrix, leverage ranking; mark C4 untested keystone, C6 double failure | The flat theory list cannot express "one experiment moves many theories"; spillover travels along cluster edges | high — drop-in JSON |
| 6 | Extend the overgeneration table + add its symmetric undergeneration counterpart — 8 cases in schema; new test_inapplicable verdict (Dualism, NCC); qualifier-burden annotations; undergeneration rows for RPT, NCC, PP | The test bites 10/12 theories but the taxonomy can't say why 2 escape; restrictiveness is currently untested | high — cases staged |
| 7 | Future research nights — pending-theory deep dives (10, Illusionism first); mine the 431-theory catalogue for a feedforward-sufficiency foil; full Nature SI extraction; human-rater replication of the day-5 audit; claim-vs-abstract audit of finding texts; analyze the now-verified Mudrik et al. 2025 review | Deferred by scope, not by value | medium — each is one night |
All re-checked 2026-08-10 against Crossref (existence + metadata + retraction). None retracted. Verification dates: each record was also verified on the day the campaign first relied on it.