Skip to main content

Backup & recovery of nodes

A Mesh Node is your organization's half of every trust decision. Its identity key is meant to stay on the host: production nodes use TPM 2.0 on Linux and Windows CNG with the Microsoft Platform Crypto Provider on Windows, so the key cannot be exported. Recovery is therefore built around three ideas:

  1. Run more than one node, so that losing one does not stop approvals.
  2. Back up configuration and customer-held material, not identity keys.
  3. Replace a lost node by enrolling a new one. Do not clone the old one.

Plan for node loss​

PracticeWhy
Run at least one more healthy node than your policy's distinct node approvals requiresIssuance continues while one node is down. With too few nodes, it fails with customer_quorum_unavailable.
Spread nodes across failure domainsA single site or cluster outage should not remove your quorum.
Assign each integration to more than one node where the adapter allows itA failed node does not block that workload path.
Leave Eligible approval nodes unchecked in profiles unless you need to pin nodesA pinned set must be edited through a policy change when a pinned node is replaced.
Enable OTLP audit exportEach node's signed audit journal is copied off the host as it is written.

What to back up​

ItemLinux locationWindows locationBack up?
Node configuration/etc/edgepki/config.jsonC:\ProgramData\Sectigo\Edge\config.jsonYes. You need it to rebuild the host the same way.
Customer-held material: local-api.crt/.key, workload-ca.pem, edge-client.crt/.key/etc/edgepki/C:\ProgramData\Sectigo\Edge\Yes, into your secret store. These are your organization's certificates and keys, not Sectigo's.
Adapter files you manage: web server configuration, Vault token file, keystore password files, OTLP client certificatesWhere you configured themWhere you configured themYes, through your normal secret and configuration management.
Node identity keyBound to the TPMBound to CNGNo. It cannot be exported, and a copy would let the copy act as the node.
State directory: audit journal, certificate slots, durable intents and outbox, cursors/var/lib/edgepki/C:\ProgramData\Sectigo\Edge\Only as a whole-host snapshot for restoring the same host. See the warning below.
Do not restore old state onto a running fleet

Sectigo Edge is built to detect rollback. Customer nodes keep every audit checkpoint they have seen and reject forks. Nonces are single-use. Cursors are verified against the journal. If you restore an older copy of a node's state directory, that node may be rejected, or may report errors such as audit_witness_fork_detected. When in doubt, enrol a new node instead.

Crash and restart behaviour​

You do not need to do anything special after a node crash, a power loss or a service restart. The node is designed to pick up where it stopped:

Interrupted duringWhat the node does on restart
Issuance, after approval but before the cloud respondedThe signed request is kept in a durable outbox and replayed idempotently. You get the same transaction, not a duplicate.
An activating rotation (NGINX, Apache, Java PKCS#12, Vault)A private, fsync-backed deployment intent records the exact transaction, CSR, private key, SANs, profile, adapter and revision. The node resumes the same rotation, without creating a new key or certificate. Abandoned intents expire within a bounded window, so keys are not kept indefinitely.
A promotion or rollbackThe adapter keeps a pending marker and reports degraded with *_recovery_pending until it can verify which certificate is serving. It never silently reports success.
HashiCorp Vault outageOnly the Vault integration degrades (vault_unavailable). It is retried on every poll and heals itself when Vault returns. Startup never contacts Vault.
OTLP exportThe signed cursor resumes from the last acknowledged event. Unsent events stay in the backlog.
Microsoft ADCS issuanceRequest-ID recovery binds an existing ADCS request back to its transaction. See Recover an ADCS request.

Check the result with sectigo-edge status and sectigo-edge integrations once the service is back.

Replace a failed or lost node​

Use this when a host is gone, its disk is unrecoverable, or you no longer trust it.

  1. Decide whether this is a security incident. If the host may have been compromised rather than simply failed, follow Incident response first. Consider pausing issuance, and revoke any certificates whose private keys lived on that host.
  2. Check the quorum. In ConsoleIncidents, check Quorum capacity. If it reads Fail closed, issuance waits until a replacement is healthy.
  3. Create a new enrollment package for the replacement host under ConsoleMesh nodes. Give it a new node name, and do not reuse the old package. See Enrollment.
  4. Install the replacement with your backed-up configuration and customer-held material. Use Linux, Windows or Helm.
  5. Move integrations. In ConsoleIntegrations, add the new node to each integration the old node served, and remove the old node. Wait until each integration shows healthy.
  6. Update pinned policies. If any workload profile pins Eligible approval nodes, or the policy references the old node in its quorum, propose a new policy version with the replacement node. Activation needs a second approver. See Policy changes & approvals.
  7. Re-establish certificates. Certificates whose keys were on the lost host must be issued again on the new node. Rotate them through the new node's adapters, and revoke the old ones if the keys may have been exposed.
  8. Confirm recovery. Check that the node is healthy in ConsoleMesh nodes, that the quorum is available, and that the audit witness quorum on ConsoleAudit evidence is met.
note

The console shows each node's status and last-seen time. A failed node keeps appearing as offline after it stops reporting. Do not reuse its name or its client certificate for the replacement. A certificate already bound to another node is rejected with node_mtls_binding_reused.

Rebuild the same host​

If the host survived but the operating system needs reinstalling, keep the TPM or CNG key provider intact and restore your configuration and material. If the key provider was cleared or replaced, for example after a TPM reset or a motherboard swap, the old identity is gone: treat the host as a new node and follow the replacement steps above.

Do not run a fresh install on a host that is still enrolled. The installer refuses with "This host is already enrolled; rerun with --upgrade and no bootstrap package". For new releases, use in-place upgrades. They change no identity, trust, token or CA material.

Recover an ADCS request​

If a Microsoft ADCS enrollment was interrupted after ADCS assigned a Request ID, an administrator can bind that existing request back to its Sectigo Edge transaction on the Windows node. The call must be authenticated with a workload identity listed in the node's microsoft_adcs_recovery_spiffe_ids setting. Any other identity is refused with adcs_recovery_forbidden.

Administrator: Windows PowerShell
.\sectigo-edge-windows-amd64.exe -ca C:\Tools\ca.pem -cert C:\Tools\recovery-svid.pem -key C:\Tools\recovery-svid-key.pem -transaction tx_9b2e71 -adcs-request-id 48211 adcs-recover

Flags must come before the command. Both -transaction and a positive -adcs-request-id are required. See Microsoft ADCS.