Research Results · August 2026

What we tested.
What happened.

This page explains the current Cairn research results in plain language. It includes what worked, what only partly worked, the important numbers and the limits of each result.

SYSTEM ONLINE
Q1 · SUPPORTED
Q2 T1 · SUPPORTED
Q2 T2 · SUPPORTED
Q3 · OPEN
How to read this page / 01

We report the result, the numbers and the limits.

A successful test tells us something useful, but it does not prove that every possible version of the problem is solved.

OUR RULE

Say what the test showed. Do not stretch it further.

If a test passes, we explain what passed. If part of a test is incomplete or fails, that stays in the record too. The goal is to make the results useful without overselling them.

Q1 Trial 001 / 02

Different model families recovered the same working state.

The first Q1 test asked whether a replacement AI model could pick up from the same preserved history without inventing a different version of the current state.

WHAT HAPPENED

The test was supported.

Two materially different model families recovered the same predefined operational state, the same unfinished commitments, the same revision relationships and the same resume point.

Neither model added unsupported claims to the authoritative state, and the two models did not disagree on the predefined assertions being scored.

WHY IT MATTERS

The system did not depend on one specific model family to remember where it was.

This supports the idea that important working state can live outside the model and be handed to a replacement model in a form it can recover reliably.

One source-tracing weakness showed up in the trial. We kept that weakness in the result instead of rewriting the test after seeing it.

WHAT THIS DOES NOT MEAN

Different models are still different.

This test was about recovering defined operational state. It does not mean two different AI models will have the same personality, reasoning style, wording or judgment in open-ended situations.

Q1 Trial 002 / 03

A smaller structured source package also worked across model families.

The follow-up asked whether the replacement model really needed a huge amount of historical text, or whether a smaller structured package could carry the important source relationships.

WHAT HAPPENED

Two complete model arms passed. One remained partial.

Claude and Grok met every required target and recovered the same required source relationships.

Gemini remains recorded as a partial result. We did not quietly remove it because the other two performed better.

WHY IT MATTERS

A replacement model did not need every old token to recover the tested relationships.

When the important history and its sources were represented directly, the completed model arms could recover the required evidence without being given the entire past as one giant block of context.

Q2 Trial 001 / 04

A million-event history did not require a million-event normal recovery.

Q2 asks whether a long-lived system can recover efficiently as its history grows, instead of rereading everything from the beginning every time it resumes.

1,000,000

Total events

The largest required Trial 001 test contained one million protected synthetic events.

10,000

Events replayed from the checkpoint

Verified checkpoint recovery rebuilt the same final state while reading only the most recent 10,000-event tail.

232

Relevant events applied

The dependency-aware path only had to apply 232 events that could actually change the final state or unfinished commitments.

0

Silent false verifications

When checkpoints, dependency data and protected history were deliberately damaged, none of the corruption tests silently returned a false verified state.

WHAT WE LEARNED

Normal recovery can start from a trusted recent point instead of replaying the whole history.

For this controlled workload, checkpoint recovery rebuilt the same final state as full replay while doing far less work.

Q2 Trial 002 / 05

Recovery time barely changed while total history grew 10×.

Trial 002 made the test stricter. We kept the recovery window fixed at exactly 10,000 events while the total lifetime history grew from 100,000 to 1,000,000 events.

MAIN RESULT● SUPPORTED
100K TOTAL HISTORY

73.733 ms

Verified checkpoint recovery with a 10,000-event recovery window.

1M TOTAL HISTORY

73.741 ms

Verified checkpoint recovery with the same 10,000-event recovery window.

FULL REPLAY · 100K

679.994 ms

Time to verify and replay the entire 100,000-event history.

FULL REPLAY · 1M

6,811.583 ms

Time to verify and replay the entire one-million-event history.

RECOVERY RATIO

1.0001

The 1M recovery time divided by the 100K recovery time. The test allowed anything up to 2.0.

OLD EVENT BODIES READ

0

Normal checkpoint recovery did not reread event bodies from before the checkpoint at either required scale.

DEPENDENCY-AWARE RECOVERY

Only a few hundred tail events actually mattered to the final state.

The dependency-aware path applied 229 relevant events at 100K total history and 232 relevant events at 1M total history while still rebuilding the same final state.

CORRUPTION TESTS

Damaged recovery data was detected instead of silently trusted.

We deliberately damaged checkpoint data, dependency data and protected history. The system rejected the bad data and either recovered safely or stopped instead of pretending the result was valid.

WHAT WE LEARNED

For this test, normal recovery depended on how much happened since the checkpoint — not on how large the entire history had become.

Full replay became roughly 10× slower when total history grew 10×. Fixed-window recovery stayed almost exactly the same.

Important limits / 06

What these tests do not prove.

The results are encouraging, but there are still clear limits to what we can say from them.

The workloads were synthetic.

They were controlled research fixtures built so the tests could be repeated exactly. Real systems may have different event patterns.

The timing came from CI runners.

The milliseconds on this page show how the approaches scaled in the test environment. They are not a promised production response time.

Q1 does not mean all models are interchangeable.

The models recovered the predefined operational state. They can still reason and behave differently.

Full-history audit still matters.

Fast checkpoint recovery is for normal resume. A full origin-to-current verification remains a different and more expensive operation.

Q2 has not tested the fixed-age design at 10 million events yet.

The required Trial 002 ladder ended at one million events. Ten million remains an extended target.

Q3 is still open.

We have not yet completed the direct comparison between structured history and brute-force long context.

What comes next / 07

Q3 is the next major comparison.

The next question is straightforward: does structured history actually outperform simply giving the model more and more old context?

NEXT QUESTION

Is structured, source-linked history more reliable than brute-force long context?

We will keep Q3 marked open until that comparison is run and scored.

CAIRN CONTINUUM · RESEARCH RESULTS

Keep the history.
Test the recovery.
Say what the results actually show.

CAIRN CONTINUUM

See where long-lived continuity could become useful in the real world.

Applications