Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Resetting Failed Recovery Qualification

Runtime-restart failure after a successful restore

Use this procedure when candidate 7c95cf54622c7e3a63d35b14db76b5e1d3dcf66c8c3af8518e3a2e9905942867 passed the Linux server, recovery backup, and recovery restore phases, then its fresh WSL2 recovery qualification failed at runtime.process_restart with:

Runtime client credential is unavailable or invalid

The v0.10.7 Managed runtime retired its still-valid bounded credential during ordinary daemon shutdown. The installed qualifier correctly stopped and restarted the daemon without logging out, but the restarted process could no longer use the same Runtime login. Release v0.10.8 preserves that credential generation across process restart and makes daemon stop wait for the old process, not only disappearance of its socket. Explicit logout, reboot, and uninstall semantics are unchanged.

Do not rerun the v0.10.7 WSL2 package and do not manually alter the restored server. The failed candidate cannot produce valid workstation or recovery evidence.

The dedicated reset accepts only the exact v0.10.7 pending-wsl2 restore:

  • candidate SHA-256 7c95cf54622c7e3a63d35b14db76b5e1d3dcf66c8c3af8518e3a2e9905942867;
  • successful linux-server.json and recovery-backup.json, with no recovery acceptance record;
  • the matching candidate binding and recovery-pending.json for restore;
  • the checksummed recovery backup and fixed 0.3.12 predecessor;
  • the cryptographically verified post-qualification Managed-secret store with its exact Deployment, recipient fingerprint, nine active entries, ten audit records, and checksum-bound safe metadata inventory; this intentionally differs from the pre-qualification backup after two failed-login cleanup histories were appended;
  • the installed v0.10.7 server, healthy v0.10.7 PostgreSQL container, active PostgreSQL and Trusted-runtime services, and declared filesystem traversal;
  • the immutable ready Datastore identity def55da2-0cf5-4fdb-8934-ec517ccaf78a, with exactly 16 sequences and 45 tables owned by its derived PostgreSQL role tre_store_def55da20cf54fdb8934ec517ccaf78a;
  • host svrltreapcc02 at 192.168.31.75, Deployment f2ef37c5-7430-468a-a439-b3ba1b0527c1, and Datastore ahri-tre-test.

It preserves the exact failed boundary beneath its candidate digest, removes only the disposable AHRI TRE installation and attempt-specific recovery state, and retains the hostname, reserved address, /data mount, Docker foundation, hosts mappings, firewall, and service identities. A mismatch is refused before mutation.

The exact replacement is candidate SHA-256 bdd0fe785ed333caf1f9c7de8b70479d5c37b5526d7ace6e28de74105c2cab48 at candidate source revision 8967fa05ec30eb13f0291f7c7a2714649a0db302. Its release components are v0.10.8 source revision a292eab3d21a02060b115927da3266f62ad982a0.

Verify it from the repository shell on the recorded WSL2 controller or an explicitly authorized Apple Silicon macOS replacement controller:

cd dist/conformance-candidate-runtime-restart-output-0.3.14
sha256sum --check --strict \
  ahri-tre-test-datastore-deployment-kit-0.3.14.tar.gz.sha256
cd ../..

Then return to the repository root and run:

./scripts/reset-failed-wsl2-recovery-conformance-wizard.sh

On macOS, the wizard first recovers the exact failed public candidate and fixed predecessor from the accepted server boundary when they are not already local. Before the destructive reset, it copies only the Runtime private leaf key and PostgreSQL server leaf key/certificate to protected paths below ~/ahri-tre-pki. It proves that the Runtime certificate and PostgreSQL public authority are unchanged between the failed and replacement candidates, then proves that both recovered keys match their replacement certificates. A mismatch stops before the server reset. These controller-recovery inputs are never placed in the repository, release archive, or evidence directory.

Approve the firewall prompt only when UFW is active with default-deny incoming, default-allow outgoing, SSH allowed, LAN HTTPS on TCP 443 allowed, and PostgreSQL TCP 5432 denied. Stop at RESET READY. Then run:

./scripts/minisforum-conformance-preparation-wizard.sh

The preparation wizard must verify the replacement checksum and stop at HOST READY. Do not reboot because the projected Secrets are ephemeral. Begin the conformance sequence again from the Linux server phase; evidence from the failed candidate is not reusable.

PostgreSQL-ownership failure during restore

Use this procedure only when installed conformance passed the Linux server and recovery-backup phases, then recovery-restore failed after committing the database because its metadata objects had the wrong owner:

ahri_tre_administrator|S|16
ahri_tre_administrator|r|45

Candidate c6e84103ca262606740d82407a79a15d63184a38013caabb86854b371623eec9 included a redundant --no-owner on its custom-format dump and actively suppressed ownership restoration with pg_restore --no-owner. The restore therefore created all 61 non-system relations—16 sequences and 45 tables— under the connecting ahri_tre_administrator role instead of retaining the immutable Datastore owner role. Installed readiness reported that the Datastore binding was not ready. Do not alter owners manually or resume recovery with that candidate. The candidate did not produce a valid restore result.

The guarded reset retires the exact failed attempt, removes its disposable installation, activates the replacement candidate, and stops before Secret preparation. It never installs a package or invokes the conformance harness.

Exact scope

The reset accepts only:

  • failed candidate SHA-256 c6e84103ca262606740d82407a79a15d63184a38013caabb86854b371623eec9;
  • replacement candidate SHA-256 7c95cf54622c7e3a63d35b14db76b5e1d3dcf66c8c3af8518e3a2e9905942867;
  • replacement source revision 1fc09bb67cdb815832ab1fb356f4a49ebc40e63c;
  • fixed 0.3.12 predecessor SHA-256 6f5555c25d96274d7772d1b46d4409bd36a42205f21c4c715de8df5179ab18ed;
  • host svrltreapcc02 at 192.168.31.75;
  • Deployment f2ef37c5-7430-468a-a439-b3ba1b0527c1; and
  • Datastore ahri-tre-test.

The filename remains ahri-tre-test-datastore-deployment-kit-0.3.14.tar.gz; its external checksum identifies the immutable candidate bytes.

The host-side guard requires the exact successful linux-server.json and recovery-backup.json records, the recovery.json no-go at restore.verify, the checksummed predecessor and backup, the restored Managed-secret and Lake contents, no recovery-qualification state, and the running managed PostgreSQL container with its durable deployment-contract bind and committed database. It also requires root:ahri-tre:0750 on /var/lib/ahri-tre, the exact declared root:ahri-tre:0750 configuration parent, root:root:0755 on /data/ahri-tre, runtime-owned 0700 Lake and scratch leaves, an active PostgreSQL service, inactive Runtime and Web services, and exactly 16 sequences plus 45 tables owned by ahri_tre_administrator. The number 142 previously reported during diagnosis was the administrator role’s PostgreSQL object ID, not an object count. A different state is refused without mutation.

What the reset preserves

The wizard preserves the three conformance records, verified recovery backup, candidate binding, and checksummed predecessor beneath the failed candidate digest. It retains the hostname, reserved address, /data mount, Docker foundation and bridge, hosts mappings, firewall configuration, and declared service identities.

It removes the active AHRI TRE package files, managed PostgreSQL container, disposable PostgreSQL data, Lake and scratch state, configuration, Managed Secrets, Injected-secret projections, Runtime state, and the attempt’s separate root-identity copy. None of the retired installation is reused.

1. Verify the replacement archive

In the ordinary WSL2 Ubuntu shell:

cd /home/kobus/repos/ahri-tre-rs/dist/conformance-candidate-postgresql-ownership-output-0.3.14
sha256sum --check --strict \
  ahri-tre-test-datastore-deployment-kit-0.3.14.tar.gz.sha256

The result must be OK, and the checksum file must contain exactly:

7c95cf54622c7e3a63d35b14db76b5e1d3dcf66c8c3af8518e3a2e9905942867  ahri-tre-test-datastore-deployment-kit-0.3.14.tar.gz

Keep the fixed predecessor archive and checksum in /home/kobus/ahri-tre-conformance/input.

2. Run the guarded reset

Leave the MinisForum unchanged and run from the repository root:

./scripts/reset-incomplete-recovery-conformance-wizard.sh

The wizard:

  1. verifies WSL2 and the exact replacement inputs;
  2. verifies the replacement checksums, source revision, recovery workflow, durable PostgreSQL deployment-contract bind, PostgreSQL ownership preservation, and shared data-parent preservation;
  3. accepts only the exact failed-restore state, verifies the backup and predecessor checksums, retires that state, and removes the bounded active installation;
  4. retires the local failed candidate pair and activates the replacement pair;
  5. proves the clean retained foundation, displays UFW, and stops at RESET READY.

Approve the firewall prompt only when UFW is active with default-deny incoming, default-allow outgoing, SSH allowed, LAN HTTPS on TCP 443 allowed, and PostgreSQL TCP 5432 denied. If any guard refuses, stop without modifying the state manually.

3. Prepare fresh Secrets

After RESET READY, run:

./scripts/minisforum-conformance-preparation-wizard.sh

The preparation wizard must use replacement SHA-256 7c95cf54622c7e3a63d35b14db76b5e1d3dcf66c8c3af8518e3a2e9905942867 and the fixed predecessor SHA-256 above. Stop when it declares HOST READY. Do not reboot because the verified /run/secrets projections are ephemeral.