Experiments
Each experiment publishes its own protocols, data and results. They are pulled straight from the open-science repository on every deploy, so what is here is what the lab has published.
- SCN fiber photometry across 14–17 days of constant darkness in 10 mice — GCaMP/GFlamp1/GFP signal, wheel-running actograms, and light-pulse phase resetting. Python port of the lab MATLAB pipeline, validated to exact match.
- There is a published taxonomy of the fourteen ways multi-agent AI systems break. We asked whether language models apply it consistently, found that they do not, and then found that our own answer was wrong: agreement scales with annotator size, from 0.173 at 2-12B to 0.598 at frontier. Underneath that sits the sharper result — the reference labels score 0.047 against themselves, while four frontier models agree with each other and not with those labels.
- Can many small open fMRI studies be merged into one shared space? Same pipeline, two naturalistic stimuli, opposite answers: alignment hurt on a non-narrative visual tone-poem (N=93) and helped on a heard narrative (19 × 8 runs) — a measured map of where pooling works, with the confounds stated
- A 7-day autonomous research campaign audits the Landscape of Consciousness dashboard: citation integrity, a Cogitate prediction-outcome ledger, a cross-theory premise graph, a blind rubric reliability audit, and the verified post-Cogitate discourse — with ranked next steps
- A research agent running an open-weight model (Kimi K3) formalized the four-valued logic FDE in Lean 4 and proved it sound, complete, and decidable — every proof checked by the kernel, every axiom reported, every wrong turn logged. The leaderboard of AI-proved math is dominated by closed models; this is the open-stack version
- Your agents write down everything they do, and almost nobody reads it back. We read ours: 2,781 sessions of real agent work, to see how AI actually breaks on your own tasks rather than on a benchmark. Most of what goes wrong turns out to be plumbing, the standard list of agent failures has no box for it, and the failures that matter most are the ones you cannot catch automatically.
- Vocalisation analysis of tūī, korimako & kākā from Xeno-Canto open data — syllable classification, geographic variation, annotation & clustering tools, and a Te Reo Manu learning game
- 14 structural critiques of the scientific paper format, an interactive evidence-overlay demo, a landscape of existing alternatives (nanopublications, executable articles, PRC, knowledge graphs), and a proposal for decoupled scientific knowledge
- Theories of consciousness catalogued, formalized with scope lines, and assessed for testability — with evidence from the 2025 Cogitate Consortium adversarial collaboration
- Community self-reports with valence — generated by Matilde