Beyond the Model: software that improved its own code, under sealed tests
Our founder's paper, released with the open-source Runesmith 1.0, shows software improving its own repair code, confirmed by two sealed tests with every result public.
This week we released a research paper, Beyond the Model, together with Runesmith 1.0, an open-source runtime for building and repairing software with AI models. Runesmith comes from AI ThinkLab, my research company. It belongs on this blog because it is built on the same principle as Veristria: proof over promises.
What Runesmith did
Runesmith lets models propose changes and lets a small deterministic kernel decide what gets applied. It changes its own code only on evidence.
In one round of its improvement loop, Runesmith picked a weakness from its own records, and a model rewrote its repair step: the part that fixes failing code. Then we tested the new version the hard way.
- Against its previous version, with the same small AI model and the same budget, the new repair step fixed 57 of 162 attempts against 35, in about half the time per repair (exact one-sided p = 0.000845).
- Against the version Runesmith shipped with, on fresh one-line bugs planted in three large public open-source projects, it fixed 42 of 168 against 17 (exact p = 0.00278).
A careful hand-written rule did better still, at 65 of 168, and the paper says so in its first paragraph. That rule is now the improvement loop's next target. The data also shows what Runesmith learned: in a post-hoc count, 123 of the 124 repairs across all three came with the faulty file in the model's view. What Runesmith kept was a better way to find the right file, and the cheaper model reused it.
Why the result holds
The rules of both tests were sealed before they ran. For the second test, the protocol and even the wording of the result were sealed before any outcome existed and publicly timestamped in the Bitcoin blockchain hours before the analysis. Six of the seven earlier sealed tests did not show their effect, and the paper reports every one of them, because the misses are what make the hits worth trusting.
Everything is public: the code, the data and scripts that recompute every test. We call this a nudge towards recursive self-improvement, and we mean exactly that: software that changed its own problem-solving code, with the change confirmed under seal.
Built with free models
Runesmith is made to get real work out of free, cheap or local models. A video composer, Runesmith Motion, was built through Runesmith with free AI models as the last author of every applied change. Anyone with a laptop and a free API key can start.
Read it, run it, check it
- The paper: doi.org/10.5281/zenodo.23196763
- Runesmith and the evidence: aithinklab.com
- The code: github.com/Veristria/runesmith
- The data and recomputations: huggingface.co/datasets/aithinklab/runesmith-evidence