Research

Do the tests.
Show the results.

Cairn is built around a hard question: can an AI system keep a trustworthy history even when the model underneath it changes? We are testing that directly instead of assuming the answer.

SYSTEM ONLINE
MODEL CHANGE · TESTED
RECOVERY · MEASURED
CORRUPTION · TESTED
RESULTS · PUBLISHED
What we are testing / 01

Five things have to work for continuity to mean anything.

The research focuses on practical questions: can the history be trusted, can the system recover, and can a different model continue from the same record?

RESEARCH PROGRAM● ACTIVE
01

Keep a trustworthy history

Test whether important events can be preserved in order and whether tampering, deletion or corruption can be detected.

02

Rebuild the current state

Test whether the system can recover what is true now, what is unfinished and what changed over time.

03

Survive model and hardware changes

Test what happens when the AI model, provider or machine underneath the system changes.

04

Measure the results

Score recovery accuracy, source tracing, unfinished work, tamper detection and model-to-model differences.

05

Test failure on purpose

Break checkpoints, damage derived memory, alter history and simulate provider loss to see whether recovery stays trustworthy.

LIMIT

Different models may still think differently

Even with the same preserved history, a replacement model may interpret some things differently. That is part of what we measure.

How we test it / 02

We deliberately create the kinds of changes that should be hard.

The system is tested under model changes, infrastructure moves, damaged memory, provider loss and altered history.

01

Replace the model

Give a different AI model the same preserved history and compare what it recovers.

02

Move the system

Change the host or infrastructure and check whether the recovered state still matches.

03

Damage derived memory

Corrupt summaries or cached memory and verify that the original history still wins.

04

Lose a provider

Test whether continuity can survive losing the model provider or the active session.

05

Tamper with history

Delete, reorder or alter protected records and verify that the change is detected.

06

Restore from backup

Recover from a verified state and compare the result with the expected baseline.

The three big questions / 03

These are the questions that can make or break the idea.

Q1 asks whether continuity survives a model change. Q2 asks whether recovery still works as history becomes very large. Q3 will compare structured history with simply giving a model a huge amount of old context.

Q1

Can a different model pick up the same work?

We test whether different model families recover the same state, unfinished commitments and source relationships from the same preserved history.

Q2

Can recovery stay fast as the history grows?

We test whether normal recovery can start from a trusted checkpoint instead of replaying everything from the beginning.

Q3

Is structured history better than just carrying more context?

This is the next major comparison. We have not marked it complete because the full test has not been run yet.

What we have learned so far / 04

Some of those questions now have real answers.

These results came from controlled synthetic tests. They are useful evidence, not promises about every future production system.

Q1 · SUPPORTED

Different model families recovered the same operational state

In Trial 001, two materially different model families recovered the same predefined state, unfinished commitments, revisions and resume point from the same preserved package.

Trial 002 then tested whether a smaller structured source package could remain portable. Claude and Grok met every required target. Gemini remains recorded as a partial result.

Trial 001 results → · Trial 002 results →

Q2 · TRIAL 001 SUPPORTED

A million-event history did not require a million-event normal recovery

At one million lifetime events, verified checkpoint recovery rebuilt the same final state while replaying only the most recent 10,000 events. The dependency-aware path needed just 232 relevant events.

View Trial 001 results →

Q2 · TRIAL 002 SUPPORTED

Recovery stayed almost unchanged while total history grew 10×

We held the recovery window at exactly 10,000 events while lifetime history grew from 100,000 to 1,000,000 events. Full replay grew from about 680 ms to about 6.8 seconds. Fixed-age recovery stayed about 73.7 ms at both sizes.

Corruption tests produced zero silent false verifications.

View Trial 002 results →

Q3 · NEXT

Structured history versus brute-force context

This test is still open. We will compare the two approaches directly instead of claiming one is better before the data exists.

See the current results →

Research notes / 05

Why this problem matters beyond Cairn.

Research Notes connect the engineering work to larger questions about AI safety, accountability and long-lived systems.

CAIRN CONTINUUM

Read the completed test results in plain language.

Research Results