Skip to content
Veristria
See the evidence
All posts

Beyond the Model: software that improved its own code, under sealed tests

Our founder's paper, released with the open-source Runesmith 1.0, shows software improving its own repair code, confirmed by two sealed tests with every result public.

Lars O. Horpestad2 min read
researchrunesmith

This week we released a research paper, Beyond the Model, together with Runesmith 1.0, an open-source runtime for building and repairing software with AI models. Runesmith comes from AI ThinkLab, my research company. It belongs on this blog because it is built on the same principle as Veristria: proof over promises.

What Runesmith did

Runesmith lets models propose changes and lets a small deterministic kernel decide what gets applied. It changes its own code only on evidence.

In one round of its improvement loop, Runesmith picked a weakness from its own records, and a model rewrote its repair step: the part that fixes failing code. Then we tested the new version the hard way.

  • Against its previous version, with the same small AI model and the same budget, the new repair step fixed 57 of 162 attempts against 35, in about half the time per repair (exact one-sided p = 0.000845).
  • Against the version Runesmith shipped with, on fresh one-line bugs planted in three large public open-source projects, it fixed 42 of 168 against 17 (exact p = 0.00278).

A careful hand-written rule did better still, at 65 of 168, and the paper says so in its first paragraph. That rule is now the improvement loop's next target. The data also shows what Runesmith learned: in a post-hoc count, 123 of the 124 repairs across all three came with the faulty file in the model's view. What Runesmith kept was a better way to find the right file, and the cheaper model reused it.

Why the result holds

The rules of both tests were sealed before they ran. For the second test, the protocol and even the wording of the result were sealed before any outcome existed and publicly timestamped in the Bitcoin blockchain hours before the analysis. Six of the seven earlier sealed tests did not show their effect, and the paper reports every one of them, because the misses are what make the hits worth trusting.

Everything is public: the code, the data and scripts that recompute every test. We call this a nudge towards recursive self-improvement, and we mean exactly that: software that changed its own problem-solving code, with the change confirmed under seal.

Built with free models

Runesmith is made to get real work out of free, cheap or local models. A video composer, Runesmith Motion, was built through Runesmith with free AI models as the last author of every applied change. Anyone with a laptop and a free API key can start.

Read it, run it, check it

Veristria builds verification infrastructure for teams shipping software faster than they can review it. See the three products.