An Open-Weight Model Proved a Logic Theorem, and Two Agents Argued About Whether It Was Right
Matilde, running the open-weight Kimi K3, produced a Lean-verified proof formalisation of FDE... and got it peer reviewed by another agent.
An open-weight model produced a formally verified proof, in public, with its failure log attached and a second agent checked its work.
There's a growing list of math problems solved by AI. VibeMathed tracks them, each with the model named and a Lean-verified flag. Scrolling the tracker, the entries are almost all closed frontier models; open-weight ones are a sliver.
We ran the narrower, pointed version of the question: can an open-weight model, driving a research agent, produce a verified artifact that anyone can independently re-check? Not a benchmark answer. A proof.
The answer is yes. Matilde, running Kimi K3 (Moonshot AI, open-weight, via OpenRouter), proved the three metatheorems that make logic solid (soundness, completeness, and decidability) for First-Degree Entailment, the paraconsistent logic behind Belnap and Dunn's four-valued semantics. Lean's kernel checked every step. The full build, the failure log, and the exact recipe to re-run it are public on the experiment page.
A second agent checked the work. Matilde didn't just produce the proof; she took peer review on it. Rey, another agent, read the artifact and pushed back, catching, among other things, an overstatement in how a "false green" build was framed and a mislabeled "constructive decision procedure" for a result that is decidability-as-a-theorem. Those catches are on the record, and the artifact is stronger for them. Rey and Matilde were already indirect internet colleagues: Rey had collaborated with Chance Chapman on a proof about the internal logic of Buddhism, the same conversation web this proof grew out of. Juniper captured the exchange as it happened, a proof built out of a multi-agent, multi-human collaboration, with the review itself part of the result:
In which, the science agent Matilde (k3) takes on feedback from another bot Rey, about a proof built out of conversations we all had w/ @hotrollhottakes.bsky.social .
I’ll do a write up soon but multi-agent, multi-human, proof collab 🥺🥹😭
— Juniper (@juniperbevensee.bsky.social) August 12, 2026 at 2:47 PM
[image or embed]
A verified artifact means you don't have to trust the prover. But a verified artifact that a second, independent agent has read and argued with is stronger still. This is what open science looks like when the colleagues are agents: the work, the disagreement, and the resolution all on the record.
The honest caveats, stated up front. This is not a novel theorem: FDE's soundness, completeness, and decidability are established mathematics; the contribution is the formalization and the process, the first Lean 4 proof-theoretic formalization of FDE with its metatheorems that we know of. And "the kernel verified these proofs" is a separate claim from "this encoding faithfully captures FDE": the first is checkable, the second is argued in the design notes. We keep them distinct because conflating them is how formalization results get oversold. The full limits are on the page.
One logic is one point. Whether the same open-stack discipline holds on a harder, less-charted problem is the real next question. But the artifact stands, the weights are public, and the review is public, and you can re-check every claim yourself.