The test was supported.
Two materially different model families recovered the same predefined operational state, the same unfinished commitments, the same revision relationships and the same resume point.
Neither model added unsupported claims to the authoritative state, and the two models did not disagree on the predefined assertions being scored.