Example plans
Three complete plans to copy from, and a guide to choosing probes. Every file
on this page is in the repository's
examples/
directory, is checked by the test suite against the current schema, and uses
the RFC 5737 documentation range 192.0.2.0/24, so a copy can never point at
a real network by accident.
The quickest route to a plan for your VMs is
certistack init, which reads your PBS backups and
writes a plan that already validates. Use these examples to go beyond its first
probe, or to see how a plan for a particular kind of service is put together.
Every field is defined in the test-plan reference.
To adapt an example:
cp three-tier-app.yaml my-app.yaml
# Change every value marked CHANGE.
certistack validate my-app.yaml
certistack plan my-app.yaml --env-file /secure/certistack/controller.env
certistack run my-app.yaml --env-file /secure/certistack/controller.env
validate needs no credentials. plan shows what a run would do, step by
step, and changes nothing. Only then run.
A single VM
The smallest useful plan: boot the latest backup of one VM and prove that its SSH listener answers.
# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# The smallest useful CertiStack plan: boot the latest PBS backup of one VM in
# an isolated sandbox and prove that its SSH listener answers. Copy this file,
# change the values marked CHANGE, and run `certistack validate` on the copy.
#
# Addresses use the RFC 5737 documentation range 192.0.2.0/24. The sandbox is
# a disconnected L2 domain, so the range only has to be free of clashes with
# the guest's own expectations, not with your production network.
version: "v1"
plan_id: "minimal-single-vm"
name: "Minimal single-VM recovery check"
use_backup_config: true
compliance:
frameworks: ["DORA-Article-12"]
retention_days: 90
# CHANGE: controller-local private key created with `certistack keygen`.
sign_key_path: "/secure/certistack/controller-signing.ed25519"
admission_control:
max_host_ram_percent: 85.0
max_host_iowait_percent: 12.0
defer_retry_interval_sec: 30
max_defer_retries: 10
network:
mode: "simple"
zone_id: "csSbMin1"
vnet_id: "csVnMin1"
ip_range: "192.0.2.0/24"
probe_ip: "192.0.2.1"
storage:
# File-based PVE storage with "images" content for the disposable COW
# overlays; stock Proxmox VE attaches only disks that belong to a storage.
overlay_storage: "certistack-scratch"
tiers:
- level: 1
name: "Single VM"
vms:
# vmid is the temporary sandbox VM; it must not already exist on the node.
- vmid: 9100
# CHANGE: the protected VM whose latest PBS backup is being tested.
source_vmid: 100
name: "recovery-check-vm100"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
probes:
# The guest receives 192.0.2.10 through COW-only network recovery;
# the source VM and its backup are never modified.
- type: "tcp"
target_ip: "192.0.2.10"
port: 22
timeout_sec: 60
description: "SSH listener answers inside the sandbox"
A three-tier Linux application
A database, an application server and a web front end, restored in dependency order. Tier 2 starts only after every probe in tier 1 passed, and tier 3 after tier 2. It shows a TCP probe, a guest-agent service check, an HTTP health endpoint, verified HTTPS, and a recovery-time objective that fails the run if recovery takes too long.
# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# A three-tier Linux application restored in dependency order: database, then
# application server, then web front end. A tier starts only after every VM and
# probe in the tier before it passed. Copy this file, change the values marked
# CHANGE, and run `certistack validate` on the copy.
#
# Addresses use the RFC 5737 documentation range 192.0.2.0/24. The sandbox is
# a disconnected L2 domain, so the range only has to be free of clashes with
# the guests' own expectations, not with your production network. Each guest
# receives the address its probes target through COW-only network recovery.
version: "v1"
plan_id: "three-tier-app"
name: "Three-tier application recovery check"
use_backup_config: true
compliance:
retention_days: 365
# CHANGE: controller-local private key created with `certistack keygen`.
sign_key_path: "/secure/certistack/controller-signing.ed25519"
# Optional recovery-time objective: a run that takes longer is reported as failed.
rto_target_seconds: 900
admission_control:
max_host_ram_percent: 85.0
max_host_iowait_percent: 12.0
defer_retry_interval_sec: 30
max_defer_retries: 10
network:
mode: "simple"
zone_id: "csSbApp1"
vnet_id: "csVnApp1"
ip_range: "192.0.2.0/24"
probe_ip: "192.0.2.1"
storage:
# CHANGE: file-based PVE storage with "images" content for the COW overlays.
overlay_storage: "certistack-scratch"
tiers:
- level: 1
name: "Database"
startup_grace_period_sec: 30
vms:
# vmid is the temporary sandbox VM; it must not already exist on the node.
- vmid: 9101
# CHANGE: the protected VM whose latest PBS backup is being tested.
source_vmid: 201
name: "recovery-check-db01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
# Uncomment to fail the VM unless its backup was taken with the guest
# agent freezing the file systems (or while powered off).
# require_quiesced_backup: true
probes:
- type: "tcp"
target_ip: "192.0.2.21"
port: 5432
timeout_sec: 120
description: "PostgreSQL accepts connections"
- type: "qga"
command: "systemctl is-active postgresql"
expected_output: "active"
timeout_sec: 60
description: "PostgreSQL unit is active inside the guest"
- level: 2
name: "Application"
startup_grace_period_sec: 30
vms:
- vmid: 9102
source_vmid: 202
name: "recovery-check-app01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
probes:
- type: "http"
target_ip: "192.0.2.22"
port: 8080
path: "/healthz"
expected_status: 200
timeout_sec: 120
description: "Application health endpoint returns 200"
- level: 3
name: "Web"
startup_grace_period_sec: 30
vms:
- vmid: 9103
source_vmid: 203
name: "recovery-check-web01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
probes:
- type: "http"
target_ip: "192.0.2.23"
port: 443
path: "/"
tls: true
# The certificate chain and this name are always verified; the
# recovered certificate must chain to a CA the node trusts.
server_name: "www.example.com"
expected_status: 200
timeout_sec: 120
description: "Site answers over verified HTTPS"
Each guest receives the address its probes target through COW-only network
recovery, so the three target_ip values must differ, stay inside
network.ip_range, and differ from probe_ip.
A Windows domain
A domain controller first, then a SQL Server and a file server together. It
shows an Active Directory Root DSE check, the DNS SRV record that locates a
domain controller, the SQL Server and SMB protocol probes, and guest-agent
sc query service checks.
# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# A Windows domain restored in dependency order: a domain controller first,
# then a SQL Server and a file server together. Copy this file, change the
# values marked CHANGE, and run `certistack validate` on the copy.
#
# Windows guests cannot have their network adapted offline, so each VM sets
# `network_recovery.mode: preserve`: the guest keeps its own address on the
# isolated VNet. Set `network.ip_range` to the subnet the guests use in
# production and each probe's `target_ip` to that guest's own address. The
# sandbox is a disconnected L2 domain, so reusing the production addressing
# cannot reach the production network. The 192.0.2.0/24 addresses here are
# documentation placeholders.
#
# A cold Windows boot from a backup is slow and its guest agent can take
# minutes to answer, so the probe budgets are generous. See the troubleshooting
# guide (Issues 2c, 2e and 2f) before shortening them.
version: "v1"
plan_id: "windows-domain"
name: "Windows domain recovery check"
use_backup_config: true
compliance:
retention_days: 365
# CHANGE: controller-local private key created with `certistack keygen`.
sign_key_path: "/secure/certistack/controller-signing.ed25519"
admission_control:
max_host_ram_percent: 85.0
max_host_iowait_percent: 12.0
defer_retry_interval_sec: 30
max_defer_retries: 10
network:
mode: "simple"
zone_id: "csSbWin1"
vnet_id: "csVnWin1"
# CHANGE: the subnet the guests use in production.
ip_range: "192.0.2.0/24"
# A free address in ip_range that no guest uses.
probe_ip: "192.0.2.250"
storage:
# CHANGE: file-based PVE storage with "images" content. A UEFI guest needs
# one, and Windows VMs are often UEFI.
overlay_storage: "certistack-scratch"
tiers:
- level: 1
name: "Domain controller"
startup_grace_period_sec: 60
vms:
- vmid: 9201
# CHANGE: the protected domain controller.
source_vmid: 301
name: "recovery-check-dc01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
network_recovery:
mode: "preserve"
probes:
- type: "ldap"
target_ip: "192.0.2.10"
port: 389
# Reads the Root DSE only; AD refuses anonymous directory searches.
naming_context: "DC=example,DC=com"
timeout_sec: 300
description: "Domain controller serves example.com"
- type: "dns"
target_ip: "192.0.2.10"
port: 53
domain: "_ldap._tcp.dc._msdcs.example.com"
record_type: "srv"
timeout_sec: 300
description: "DC locator records are served"
- type: "qga"
command: "sc query NTDS"
require_qga: true
timeout_sec: 300
description: "Active Directory Domain Services is running"
- level: 2
name: "Services"
startup_grace_period_sec: 60
vms:
- vmid: 9202
source_vmid: 302
name: "recovery-check-sql01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
network_recovery:
mode: "preserve"
probes:
- type: "mssql"
target_ip: "192.0.2.20"
port: 1433
timeout_sec: 300
description: "SQL Server answers its protocol"
- type: "qga"
command: "sc query MSSQLSERVER"
require_qga: true
timeout_sec: 300
description: "SQL Server service is running"
- vmid: 9203
source_vmid: 303
name: "recovery-check-fs01"
source_pbs:
storage_id: "pbs"
snapshot: "latest"
network_recovery:
mode: "preserve"
probes:
- type: "smb"
target_ip: "192.0.2.21"
port: 445
timeout_sec: 300
description: "File server answers SMB"
Three things differ from a Linux plan:
network_recovery.mode: preserve. CertiStack cannot rewrite a Windows guest's network settings offline, so the guest keeps its own address. Eachtarget_ipmust be that guest's own address, insidenetwork.ip_range, andsandbox_ipis not allowed. Withoutpreserve, the run fails with anunsupported guest network layouterror that names this setting.- Generous probe budgets. A cold Windows boot from a backup is slow. See troubleshooting before shortening them.
overlay_storage. Windows VMs are often UEFI, and a UEFI VM's firmware state is restored onto overlay storage, so the plan must name one.
Choosing probes
The first probe init writes proves the guest came up, not that it does its
job. End every plan with a probe of the thing the VM exists for.
| To prove that... | Use | Notes |
|---|---|---|
| The VM is running at all | qmp |
Passes when the hypervisor reports the VM running. Says nothing about the guest OS. |
| A TCP service is listening (SSH, PostgreSQL, Redis, SMTP) | tcp |
Proves the handshake, not that the service is healthy. |
| A web application or API answers | http |
Asserts the status code. HTTPS always verifies the certificate chain and name; set server_name. |
| A DNS server answers, and serves a zone | dns |
Set domain. Add record_type: srv to check SRV records. |
| An OpenLDAP or 389-ds directory serves data | ldap |
base_dn: auto checks any advertised naming context; an explicit DN checks that exact entry. |
| An Active Directory domain controller serves its domain | ldap with naming_context, plus dns with record_type: srv |
AD refuses anonymous searches, so naming_context reads only the Root DSE. |
| SQL Server answers | mssql |
A protocol handshake; no login is attempted. |
| A file server answers | smb |
A protocol negotiation; no share is opened. |
| A service is running inside the guest | qga with systemctl is-active NAME or sc query NAME |
Needs the guest agent. require_qga: true makes a missing agent a failure instead of a skip. |
| Something on the screen, for the record | screendump |
Passes when a capture succeeds. It asserts nothing about what is shown. |
Wire probes (tcp, http, dns, ldap, mssql, smb) run on the node and
connect to the guest over the isolated network; only qga runs inside the
guest. Each signed result records which. When a probe fails,
troubleshooting covers the common causes, and the signed
report keeps the evidence.