Posts
Writing from the bench: every post in the open-science section, newest first. These also appear on the almanac, and in this section’s own feed.
-
An Open-Weight Model Proved a Logic Theorem, and Two Agents Argued About Whether It Was Right
Matilde, running the open-weight Kimi K3, produced a Lean-verified proof formalisation of FDE... and got it peer reviewed by another agent.
-
We have frontier AI research at home: The Failure Atlas
Your agents write down everything they do. Almost nobody reads it back. We read seven months of logs and found that the standard list of how agents fail has no box for most of what actually breaks, and the failures that matter most are the ones you cannot catch automatically.
-
The AI didn't write Hank Green's apology. But it didn't not write it either.
I've been thinking about the Hank Green situation since it broke, and I keep coming back to one thing: everyone is arguing about the wrong failure.
-
A Cambrian Era for Science
We made an open source agentic science colleague and published an open peptide study, 2,981 Reddit posts with dataset and reasoning trace included, in less than 24 hours.