Skip to main content

Monitoring & alerts

Sectigo Edge gives you signals from two places: the console, which shows the tenant-wide view, and each Mesh Node, which exposes local health, Prometheus metrics, the signed trust-event feed and an optional OTLP audit stream. Sectigo Edge has no built-in paging. Connect the node endpoints and event feeds to the monitoring and on-call tools you already use.

What to monitor​

SignalSourceAlert whenSeverity
Node healthGET /healthz on each nodestatus is not ok, or the endpoint stops answeringHigh
Node status in the consoleConsoleMesh nodesA node is degraded, offline or disabledHigh
Approval quorumConsoleIncidents, Quorum capacityFail closed: too few healthy nodes to approve issuanceCritical
Signed policy syncedgepki_node_policy_sync_total{result="failure"}Failures keep increasing, or policy_last_synced_at in /healthz stops advancingHigh
Integration healthsectigo-edge integrations, or ConsoleIntegrationsAny assigned integration is degraded, or stuck in pending_nodeHigh
Certificate expiryConsoleTrust health, Expiry exposureAbove 0, which means active certificates expire within 30 daysMedium
Rotation progressedgepki_node_rotations_promoted_totalNo increase during an expected renewal periodMedium
Issuance pausedAudit issuance.paused, or trust event incident.issuance_pausedAlwaysCritical
Resume attemptsAudit break_glass.resume_requested, break_glass.resume_expired, issuance.resumedAlwaysHigh
Revocation publicationConsoleRevocationA CRL is expired or pending, or OCSP responses need attentionHigh
Audit integrityConsoleAudit evidenceAudit chain could not be verifiedCritical
Audit witness quorumConsoleAudit evidenceThe witness status is not healthy, so evidence export would waitMedium
OTLP audit exportNode metricsBacklog growing, failures repeating, or the cursor not advancingHigh
Trust-event cursorYour subscriberHTTP 410 (cursor expired)High

Console indicators​

Trust health page with metrics for managed identities, healthy nodes, expiry exposure and policy version
Trust health summarizes node health, expiry exposure and the active policy version.

The status label in the top bar summarizes the tenant:

LabelMeaning
trust path healthyAll nodes are healthy and trust status is healthy.
Degraded · n/m nodes healthySome enrolled nodes are not healthy.
DegradedEvery node is healthy, but tenant trust status is degraded.
No Mesh Nodes connectedNo node is enrolled, so customer approval is impossible.
issuance pausedThe kill switch is active. See Incident response.

The Healthy nodes metric reads, for example, 3/3 with All enrolled nodes reporting healthy, or tells you how many nodes are not healthy. The ConsoleMesh nodes table shows each node's version, capabilities, last seen time and status.

Node health endpoint​

Each Mesh Node answers GET /healthz on its local API. The sectigo-edge status command prints the same output:

mesh-node-01 (bash)
sectigo-edge -ca /etc/sectigo-edge/ca.pem \
  -cert /run/spire/svid.pem -key /run/spire/svid-key.pem status
{
"status": "degraded",
"node_id": "node_7f3c9a",
"policy_version": 12,
"policy_sha256": "pV0f3m...",
"policy_last_synced_at": "2026-10-01T14:05:02Z",
"audit_head": "Jx2cA9...",
"audit_sequence": 18342,
"integrations": [
  { "kind": "nginx", "installed": true, "health": "healthy" },
  { "kind": "otlp_audit", "enabled": true, "health": "degraded",
    "error_code": "otlp_audit_transport_failed", "exporter_id": "primary-siem",
    "endpoint_host": "otel-collector.example.com", "cursor": 18011,
    "source_head": 18342, "backlog": 331 }
],
"dependency_discovery": { "enabled": false, "configured_probes": 0,
  "successful": 0, "failed": 0, "queued": 0 }
}

status becomes degraded when the OTLP exporter is failing. It also becomes degraded when dependency discovery is enabled and either its most recent probe failed or observations are waiting to be sent. A degraded node still answers and still participates in trust decisions. Use the per-integration health and error_code fields to find the cause.

Prometheus metrics​

Scrape GET /metrics on each node's local API. All metrics use the edgepki_node_ prefix.

MetricTypeUse it to
edgepki_node_requests_totalcounterTrack issuance requests evaluated by this node
edgepki_node_policy_decisions_total{decision="allow"|"deny"}counterSpot a jump in local denials
edgepki_node_certificates_issued_totalcounterConfirm certificates are coming back from the signer
edgepki_node_outbox_queued_totalcounterSee approved transactions queued after transport failures
edgepki_node_rotations_promoted_totalcounterConfirm health-gated promotions are happening
edgepki_node_policy_sync_total{result="success"|"failure"}counterAlert on signed policy sync failures
edgepki_node_registration_sync_totalcounterTrack service-registration sync
edgepki_node_trust_event_batches_total, edgepki_node_trust_events_delivered_totalcounterTrack trust-event feed usage
edgepki_node_dependency_probes_total, edgepki_node_dependency_observations_totalcounterTrack dependency discovery
edgepki_node_otlp_audit_exports_total{result="success"|"failure"}counterOTLP export results (only when OTLP is enabled)
edgepki_node_otlp_audit_cursorgaugeLast sequence acknowledged by the collector
edgepki_node_otlp_audit_source_headgaugeNewest local audit sequence
edgepki_node_otlp_audit_backloggaugeEvents not yet acknowledged

Example alert rules:

groups:
- name: sectigo-edge-node
rules:
- alert: SectigoEdgePolicySyncFailing
expr: increase(edgepki_node_policy_sync_total{result="failure"}[15m]) > 0
and increase(edgepki_node_policy_sync_total{result="success"}[15m]) == 0
for: 15m
labels: { severity: page }
- alert: SectigoEdgeOtlpBacklogGrowing
expr: deriv(edgepki_node_otlp_audit_backlog[30m]) > 0
for: 30m
labels: { severity: ticket }
- alert: SectigoEdgeOtlpCursorStalled
expr: changes(edgepki_node_otlp_audit_cursor[30m]) == 0
and changes(edgepki_node_otlp_audit_source_head[30m]) > 0
labels: { severity: page }

Adjust the thresholds and windows to fit your estate.

Integration status codes​

Each integration in ConsoleIntegrations has one of four states:

StateMeaning
disabledNo authority or workload path is active.
pending_nodeDesired state, or a test, is waiting for the node to acknowledge it.
healthyThe node applied and checked the exact current revision.
degradedThe node rejected the adapter or could not operate it. The error_code says why.

The common error_code values are listed in Troubleshooting → Integrations.

note

The Sectigo SCM integration runs in the cloud and has no live connectivity check. When you save or test it, it shows degraded with scm_connectivity_not_verified, and the test message reads Test recorded, but connectivity was not verified. This is expected. Confirm SCM connectivity by issuing a test certificate.

OTLP audit export​

WatchWhy
edgepki_node_otlp_audit_backlog growingThe collector is behind or unreachable. Nothing is lost, but evidence is delayed.
edgepki_node_otlp_audit_exports_total{result="failure"} repeatingLook at the exporter's error_code in /healthz.
Cursor flat while the source head growsExport has stalled.
Client certificate expiryAn expired client identity shows as otlp_audit_client_identity_expired.

Exporter error codes include otlp_audit_transport_failed, otlp_audit_collector_rejected, otlp_audit_partial_success, otlp_audit_response_invalid, otlp_audit_client_identity_expired and otlp_audit_cursor_commit_failed. Setup is described in Audit & evidence export.

Trust events​

Trust events page explaining the signed event flow, events eligible for projection and SDK and CLI subscriptions
Trust events lists the facts that can be distributed to subscribers and shows how subscribers verify them.

Signed trust events let your automation react quickly without polling the console. The topics are capability.observed, capability.expired, service.claimed, identity.rotated, trust_bundle.updated, policy.version.available, certificate.staged, certificate.promoted and incident.issuance_paused.

Follow them from a node with the CLI:

observer-host (bash)
sectigo-edge \
  -url https://127.0.0.1:9443 \
  -ca /etc/sectigo-edge/ca.pem \
  -cert /run/spire/svid.pem \
  -key /run/spire/svid-key.pem \
  -tenant acme-corp \
  -node-spki /etc/sectigo-edge/node-public.pem \
  -node-spki-sha256 "$SECTIGO_EDGE_NODE_SPKI_SHA256" \
  -event-cursor-file /var/lib/sectigo-edge/event-cursors.json \
  -event-topics certificate.promoted,incident.issuance_paused \
  -event-wait 20s -watch events

Flags must come before the command. Watch mode writes one verified, signed event per line. You can pipe that output into your alerting tool.

warning

Events are notifications, not authority. An event can trigger a refresh, but it can never approve, activate, revoke or resume anything. When a cursor expires, the command stops and the API returns HTTP 410. Alert on this and resynchronize deliberately. Never overwrite the cursor file automatically.

The local event endpoint is disabled unless you allow specific observer identities in the node configuration (trust_event_subscriber_spiffe_prefixes). See Node configuration.