Keep a trustworthy history
Test whether important events can be preserved in order and whether tampering, deletion or corruption can be detected.
Cairn is built around a hard question: can an AI system keep a trustworthy history even when the model underneath it changes? We are testing that directly instead of assuming the answer.
The research focuses on practical questions: can the history be trusted, can the system recover, and can a different model continue from the same record?
Test whether important events can be preserved in order and whether tampering, deletion or corruption can be detected.
Test whether the system can recover what is true now, what is unfinished and what changed over time.
Test what happens when the AI model, provider or machine underneath the system changes.
Score recovery accuracy, source tracing, unfinished work, tamper detection and model-to-model differences.
Break checkpoints, damage derived memory, alter history and simulate provider loss to see whether recovery stays trustworthy.
Even with the same preserved history, a replacement model may interpret some things differently. That is part of what we measure.
The system is tested under model changes, infrastructure moves, damaged memory, provider loss and altered history.
Give a different AI model the same preserved history and compare what it recovers.
Change the host or infrastructure and check whether the recovered state still matches.
Corrupt summaries or cached memory and verify that the original history still wins.
Test whether continuity can survive losing the model provider or the active session.
Delete, reorder or alter protected records and verify that the change is detected.
Recover from a verified state and compare the result with the expected baseline.
Q1 asks whether continuity survives a model change. Q2 asks whether recovery still works as history becomes very large. Q3 will compare structured history with simply giving a model a huge amount of old context.
We test whether different model families recover the same state, unfinished commitments and source relationships from the same preserved history.
We test whether normal recovery can start from a trusted checkpoint instead of replaying everything from the beginning.
This is the next major comparison. We have not marked it complete because the full test has not been run yet.
These results came from controlled synthetic tests. They are useful evidence, not promises about every future production system.
In Trial 001, two materially different model families recovered the same predefined state, unfinished commitments, revisions and resume point from the same preserved package.
Trial 002 then tested whether a smaller structured source package could remain portable. Claude and Grok met every required target. Gemini remains recorded as a partial result.
At one million lifetime events, verified checkpoint recovery rebuilt the same final state while replaying only the most recent 10,000 events. The dependency-aware path needed just 232 relevant events.
We held the recovery window at exactly 10,000 events while lifetime history grew from 100,000 to 1,000,000 events. Full replay grew from about 680 ms to about 6.8 seconds. Fixed-age recovery stayed about 73.7 ms at both sizes.
Corruption tests produced zero silent false verifications.
This test is still open. We will compare the two approaches directly instead of claiming one is better before the data exists.
Research Notes connect the engineering work to larger questions about AI safety, accountability and long-lived systems.
Bill Gates recently wrote about the difficult choices ahead as AI becomes more capable and autonomous. This note looks at one part of that challenge: keeping a trustworthy history as models and infrastructure change.