Skip to content

Declarative Test Plan Specification & Reference

CertiStack test plans are written in YAML (version: "v1"). Sections 1–4 describe the recommended plan layout, probe definitions, and the literal-plan rule. Section 5 lists every key the parser accepts, treats as a deprecated alias, or deliberately rejects.


1. Top-Level Schema Structure

version: "v1"
plan_id: "string (alphanumeric, hyphens, underscores)"
name: "Human readable plan title"
# Optional: inherit safe compute settings and all backed-up disks from
# each resolved PBS snapshot. Network and PVE ownership settings stay isolated.
use_backup_config: true

compliance:
  frameworks: [ "DORA-Article-12", "SOC2-Type2-CC7.4", "HIPAA-164.308-A7", "ISO27001-A8.13" ]  # optional
  retention_days: int (0 to 36500; optional)
  sign_key_path: "string (/path/to/private.key)"

admission_control:
  max_host_ram_percent: float (> 0 and <= 100)
  max_host_iowait_percent: float (> 0 and <= 100)
  defer_retry_interval_sec: int (> 0)
  max_defer_retries: int (> 0)

network:
  mode: "simple" | "vxlan"
  zone_id: "string"
  vnet_id: "string"
  tag: int (VXLAN VNI 1..16777215; required for vxlan)
  peers: "comma-separated peer IPs (required for vxlan)"
  ip_range: "CIDR string (e.g. 192.0.2.0/24)"
  probe_ip: "host address inside ip_range (e.g. 192.0.2.1)"

tiers:
  - level: int (>= 1, executed in ascending order)
    name: "string"
    startup_grace_period_sec: int
    boot_timeout_sec: int (>= 0; default 60; applies to PVE start and QMP readiness)
    vms:
      - vmid: int (> 0, unique across entire plan; temporary test VM)
        source_vmid: int (> 0; protected source VM)
        name: "string"
        source_pbs:
          # `pbs` selects this run's configured PBS endpoint. Other selectors
          # are rejected instead of silently routing a VM to a different backup.
          storage_id: "pbs"
          snapshot: "latest" | "pinned PBS snapshot path"
          drives:
            - slot: "scsi0"
              archive: "drive-scsi0"
        hardware_overrides:
          cores: int
          sockets: int
          memory_mb: int
          balloon_mb: int
          cpu_weight: int
        probes:
          - type: "tcp" | "http" | "dns" | "ldap" | "mssql" | "smb" | "qga"
            timeout_sec: int (> 0; total probe readiness budget)
            # Optional. 0 keeps automatic polling until timeout; 1-10 limits
            # the probe to that many additional executions.
            retries: int (0 to 10)
            ...

When use_backup_config: true, CertiStack discovers the backed-up disk layout and safe compute/firmware settings from qemu-server.conf. It also preserves the source net0 MAC and NIC model while attaching that NIC only to the isolated VNet; this keeps MAC- and driver-bound guest network profiles working without replaying a production bridge, firewall setting, VLAN, gateway, or source PVE tags.


2. Configuration Field Definitions

compliance

Controls audit metadata and non-repudiation signing.

Field Type Required Description
frameworks []string No Target regulatory standards recorded in the report; the list may be empty. Canonical items: DORA-Article-12, SOC2-Type2-CC7.4, HIPAA-164.308-A7, ISO27001-A8.13; section 5 lists the short forms and variants that are also accepted. A framework label is not a certification.
retention_days int No Audit record retention the report states (e.g., 2555 for 7-year financial retention), from 0 (the default) to 36500. It is recorded in the signed report for whoever keeps it; CertiStack never prunes or deletes a report.
sign_key_path string Yes Controller-local path to the Ed25519 private key used to sign the final audit certificate. The path is never sent to the PVE node worker. --key or CERTISTACK_SIGN_KEY overrides it at run time, but validate still requires it in the plan.

admission_control

Protects production workloads running on the hypervisor host.

The first four keys are required: a plan that omits one fails validation. The values in the examples are the ones shown below. max_concurrent_boots and boot_settle_sec are optional and have the defaults shown.

Field Type Required / default Description
max_host_ram_percent float required; examples use 85.0 Maximum host RAM utilization allowed before starting a test.
max_host_iowait_percent float required; examples use 12.0 Maximum host CPU I/O wait % allowed before starting a test.
defer_retry_interval_sec int required; examples use 30 Seconds to sleep between admission re-evaluations.
max_defer_retries int required; examples use 10 Maximum retry attempts before marking the test run as throttled/failed.
max_concurrent_boots int optional, 0 (unlimited) How many sandbox VMs may boot at once on the node. Before starting a VM, the run waits while this many VMs started within boot_settle_sec.
boot_settle_sec int optional, 180 How long a started sandbox VM counts as booting for max_concurrent_boots.

A tier starts its VMs one after another, as soon as each is prepared, so their first minutes of boot overlap. That is when each guest writes most to its new overlays. On a node whose overlay storage shares a disk with the cluster file system's database (the run warns about it), several overlapping boots can delay corosync and cost the node its cluster membership. max_concurrent_boots: 1 or 2 spaces them out, across tiers too. The wait is logged, counted in the run's admission time, and shown by certistack plan before each start it may hold.

In release mode, each VM's pre-boot integrity scan reads all of its disks before it starts, so the starts already come about one scan apart: in the lab, a tier of three VMs with 10-minute scans never waited. The limit holds starts back mostly for VMs with small disks, whose scans are short, and in test mode, which has no scan.

Before creating any sandbox resource, CertiStack also reserves the declared maximum memory for the largest concurrent recovery tier plus 256 MiB per sandbox VM for host overhead. It combines that reservation with current host usage and the RAM percentage cap. Ballooning is not counted as guaranteed free capacity. This fails safely before PBS mapping when a full or OS-wide matrix would overcommit the node.

network

Configures isolated Proxmox Software-Defined Network (SDN) sandboxes.

Field Type Options Description
mode string simple, vxlan simple for single-node air-gapped bridges; vxlan for multi-node clusters.
lifecycle string ephemeral (default), preprovisioned ephemeral makes CertiStack create and remove the zone/VNet. preprovisioned uses a platform-owned sandbox only after read-only validation of its zone, VNet relationship, subnet, no-gateway, and no-SNAT invariants.
zone_id string — Unique PVE SDN identifier (e.g. csZone01; letters and digits only, starting with a letter).
vnet_id string — Unique PVE SDN identifier (e.g. csVnet01; letters and digits only, starting with a letter).
tag int 1..16777215 for vxlan VXLAN VNI assigned to the VNet. It must be explicit and unused by every existing PVE VNet; simple networks may omit it.
peers string comma-separated IP addresses for vxlan Required VXLAN zone peer list. CertiStack canonicalizes the addresses and sends them as the PVE zone peers field; simple networks must omit it.
ip_range string CIDR Isolated subnet CIDR block (e.g. 192.0.2.0/24). Strictly no gateway or SNAT is provisioned. It must not overlap loopback, unspecified, link-local, multicast, or reserved space, which the worker host would answer itself.
probe_ip string IP address in ip_range Required temporary address assigned by the node worker to the PVE VNet bridge for Stage 2 probes. It must not be the subnet network/broadcast address or a guest target. The worker requires per-interface IPv4 forwarding to be disabled and installs scoped host-input/forwarding containment before assignment. The address adds only the connected subnet route and is removed during teardown/recovery; it is never a gateway or SNAT address. Wire probes (tcp, http, dns, ldap, mssql, smb) are sent from this address and bound to the VNet interface, so neither the host routing table nor a host listener can answer them.

The IDs may remain stable across scheduled runs. With the default lifecycle: ephemeral, a successful CertiStack run removes the resources it owns. Before it creates anything, CertiStack reads the exact zone and VNet IDs. If either already exists, it fails closed without reusing, modifying, or deleting it. This prevents a validation plan from adopting a customer or a previously interrupted sandbox. Review the matching PVE resource or reconcile the interrupted run, then retry the same plan; do not work around the condition by repointing a production configuration.

Applying SDN in PVE is all or nothing, on every node, so CertiStack creates and removes its zone and VNet under PVE's SDN configuration lock and applies only its own change. PVE refuses the lock while the cluster's SDN has changes nobody applied, and the run is then refused before it creates anything (troubleshooting Issue 3d). plan reports it too. The lock is held for seconds, around each create-and-apply and remove-and-apply, and needs SDN.Allocate on /sdn, which the documented role grants.

Use lifecycle: preprovisioned when the platform operator deliberately provisions one dedicated, air-gapped sandbox ahead of time. CertiStack performs only read-only SDN API requests for that sandbox and retains it at teardown; it still adds and removes the exact host-side probe address and destroys every test VM, overlay, and PBS map it owns. A report records the validated-and-retained sandbox rather than falsely claiming it was purged.

PVE's ifupdown2 turns IPv4 forwarding on for every SDN bridge at each network reload. CertiStack refuses a sandbox VNet that forwards, because that would let the host route between the sandbox and other networks. For an ephemeral sandbox, the bridge is the run's own: the run waits for the node's network reload to bring it up and turns forwarding off itself. The containment table it adds with the probe address drops forwarded packets as well, even if a later network reload turns forwarding back on. A pre-provisioned VNet is the platform's, so the platform keeps forwarding off on it, for example with an /etc/network/if-up.d/ hook that writes 0 to /proc/sys/net/ipv4/conf/<vnet>/forwarding for that VNet on every node that may run a validation. A run otherwise refuses it before adding the probe address.

tiers & Dependency DAG

Tiers define the sequential boot order of your application architecture. Tiers are sorted by level in ascending order ($1 \to 2 \to 3$).

  • CertiStack starts each planned VM in Tier $N$, waits for startup_grace_period_sec, and evaluates Stage 2 & Stage 3 probes for every VM that reached Stage 1.
  • A failure does not hide its independent peers: CertiStack collects every current-tier outcome and a failure framebuffer per affected VM in the signed report.
  • Only after all VMs and probes in Tier $N$ pass does Tier $N+1$ begin execution. A failed tier blocks dependent tiers.

Snapshot selection

Set each VM's source_vmid and source_pbs.snapshot: "latest" for normal scheduled validation. Before any SDN or VM resource is created, CertiStack queries that VM's PBS backup group, selects the highest backup-time, and replaces latest with the exact immutable snapshot path. That resolved path is included in the report. A full snapshot path remains supported only when an operator intentionally needs to replay a historical recovery test.

Source disk layout

Use source_pbs.drives for every backed-up disk. Each entry keeps its source PVE slot on any bus (scsi0–scsi30, virtio0–virtio15, sata0–sata5, ide0–ide3) and names the PBS archive. Do not include Cloud-Init CDs or other non-backed-up media. source_pbs.drive remains accepted for existing single-disk plans; its slot is the one its archive names (drive-virtio0 is virtio0), else scsi0.

Every run reads the snapshot's qemu-server.conf, whatever the plan inherits, and applies what a restore cannot do without:

  • Boot-critical hardware. Firmware (see below), machine type, SCSI controller and guest OS type come from the backup unless the plan sets them, because a guest may not boot on PVE's defaults: a Windows guest installed on an LSI controller has no driver for virtio-scsi.
  • Machine version as a restore resolves it. Proxmox VE pins a new Windows VM's machine version on create (q35 becomes, for example, pc-q35-11.0+pve2), which a restore does not: qmrestore keeps the backup's setting, and PVE resolves an unversioned one at every start. The sandbox VM gets the backup's setting back after create, and the report records the version it actually ran as restored_hardware.running_machine. For a Windows guest, PVE starts an unversioned machine type on the version of the QEMU that created the VM (its meta line), or on 5.1 when that QEMU is older than 9.1. PVE writes the sandbox VM's meta itself and refuses to change it, so when the backup's version differs from the node's latest, the sandbox VM is pinned explicitly to the version a restore would start on.
  • SMBIOS identity as a restore keeps it. The sandbox VM gets the backup's smbios1 (its UUID and any manufacturer, product, serial, SKU and family strings), as qmrestore keeps it; a guest can tell a new SMBIOS UUID from its own. The VM generation ID is new, as after any restore. The report records the UUID as restored_hardware.smbios_uuid.
  • No missing drives. A restore keeps the source's CD-ROM drives (a cloud-init drive, an ISO), so the guest finds the same controllers. The sandbox VM gets them empty, so no configuration drive is regenerated with the source's network settings. The report lists them as restored_hardware.empty_cdroms.
  • No new devices. A source without a memory balloon device (balloon: 0) gets none in the sandbox either, unless the plan sets hardware_overrides.balloon_mb; a device the guest never had is new hardware to it. The report records this as restored_hardware.no_balloon.
  • Boot disk. The sandbox VM boots the first restored disk in the source's boot order (boot: order=..., or the legacy bootdisk). Without one it boots the disk PVE would pick: the first by bus (ide, scsi, virtio, sata) and index. The report records it as restored_hardware.boot_disk.
  • Disks the backup has. A plan that lists a disk the backup does not contain is refused.
  • Disks the plan leaves out. A plan may restore only some of the backed-up disks, for example the boot disk alone. The disks it leaves out are listed in the VM's omitted_source_disks in the signed report, and the run warns, so a partial restore is never presented as the whole VM. Each archive must go to the slot it was backed up from, and only once.
  • Disks no backup contains. Disks the source excludes from backups (backup=0) are listed as excluded_source_disks: no restore can include them, so the evidence says the VM was never whole in PBS.
  • Protection. source_crypt_mode is the weakest PBS crypt mode among the VM's disk archives (encrypt, sign-only or none), and source_signature_verified is true when the configured key verified the snapshot manifest's signature.
  • Consistency. source_consistency says how consistent the backed-up disks were, from the backup's own log (client.log.blob, which vzdump stores with the snapshot):
  • quiesced: the guest agent froze the file systems (VSS on Windows);
  • crash-consistent: a running VM was backed up without a freeze that held (the agent was missing or not running, the freeze was refused, for example by SELinux, or Windows released it before the snapshot);
  • powered-off: the VM was not running, or was shut down for the backup;
  • unknown: the log could not be read. source_consistency_detail gives the reason in the log's words, and a crash-consistent backup is flagged in the run's progress. It fails a VM only when the plan sets require_quiesced_backup: true on it: then anything but quiesced or powered-off, unknown included, fails that VM before it is restored, and doctor fails it too. The log is read through the PBS API with the configured repository, credentials and certificate fingerprint, because proxmox-backup-client cannot fetch a file outside the snapshot manifest. It is not covered by the manifest's signature, and an encrypted backup's log is reported as unknown.
source_pbs:
  storage_id: "pbs"
  snapshot: "latest"
  drives:
    - slot: "scsi0"
      archive: "drive-scsi0"
    - slot: "scsi1"
      archive: "drive-scsi1"

Backup configuration inheritance

Set use_backup_config: true to have CertiStack read the exact resolved snapshot's qemu-server.conf and discover every disk image archive (drive-scsiN, drive-virtioN, drive-sataN, drive-ideN). The inherited allowlist is CPU cores/sockets, memory/ballooning, CPU weight, machine type, guest OS type, firmware, and SCSI controller. Explicit plan values override these individual inherited fields.

CertiStack never replays source network interfaces, PVE tags, hooks, startup rules, host PCI/USB passthrough, or arbitrary QEMU arguments. Source tags are preserved as attestation evidence only. It creates a fresh isolated NIC and applies only CertiStack ownership tags.

UEFI and TPM firmware state

A UEFI source VM (bios: ovmf) keeps its firmware state in two small disks: the UEFI variable store (efidisk0: boot entries and Secure Boot keys) and, with a virtual TPM, the TPM state (tpmstate0). CertiStack restores both with every plan, whether or not it sets use_backup_config, because a UEFI VM restored without them, or under SeaBIOS, is a different machine:

  • Each image is restored whole from PBS into a new raw volume on storage.overlay_storage, which UEFI VMs therefore require. The firmware writes its variable store and PVE keeps TPM state only in raw volumes, so a read-only map cannot be used. The protected backup is never written.
  • The PBS client checks every chunk while restoring. The report lists each image under source_disks with its archive and SHA-256, and under restored_hardware.firmware_disks with its options, for example efidisk0 efitype=4m,pre-enrolled-keys=1 and tpmstate0 version=v2.0. The images are always fully restored, test mode included.
  • Only the options that describe the image are carried over: efitype, pre-enrolled-keys and ms-cert for efidisk0, and version for tpmstate0.
  • A plan that leaves bios empty restores a UEFI backup under OVMF. A plan that sets bios: seabios for a UEFI backup is refused.
  • A firmware disk that was excluded from the backup (backup=0) fails closed, because the source firmware state cannot be reproduced.
  • The copies hold the source's firmware secrets: Secure Boot variables, and TPM-sealed keys such as a BitLocker protector. They are 0600, journaled like every overlay, and removed with the sandbox VM, including by crash recovery.
  • Data disk overlays carry backup=0, but PVE has no such flag for efidisk0 and tpmstate0. A cluster backup job covering all VMs that runs while a validation is in progress therefore copies the sandbox VM's firmware state (its TPM secrets included) into a new backup group, under that job's retention and permissions. Exclude the sandbox VMIDs from such jobs (--exclude, or pool-based selection) or schedule validations outside the backup window.

Copy-before-boot

By default (copy_before_boot: auto), when a tier's disks exceed the cache budget where reading directly from PBS would cause parts of the guest boot to miss service start deadlines, CertiStack copies the disks to local overlay storage during the integrity scan so the sandbox boots from local copies as a full restore does.

  • auto (default): copies when disks would not all stay cached until boot, provided overlay storage is on a disk separate from the cluster database and has sufficient free space for the disks plus a 5 GiB reserve.
  • always: unconditionally copies disks or fails the VM before boot.
  • never: boots directly from the read-only PBS mapped devices through COW overlays.

storage.max_copy_gib (default: 0, uncapped) limits the cumulative size of copies made in a single run. Reports record copied_before_boot and each disk's boot_source.

COW-only guest network recovery

Some static guests bind their production connection to a source interface name or PCI path. That binding can change in an isolated recovery VM even when its MAC is preserved. For a VM with Stage 2 wire probes, CertiStack automatically creates one temporary connection profile on the boot disk's disposable COW overlay before boot. You normally need this stanza only for an optional address, legacy interface name, or failure diagnostics:

network_recovery:
  mode: auto
  # Optional when all Stage 2 probes already target one address.
  sandbox_ip: "192.0.2.10"
  # Optional compatibility fallback for a legacy guest profile bound to a
  # predictable Linux device name; writes a MAC-bound COW-only .link file.
  interface_name: "eth0"
  # Optional: preserve useful guest-side diagnostics when a wire probe fails.
  capture_diagnostics: true

auto detects legacy RHEL/Rocky ifcfg, NetworkManager, Netplan, systemd-networkd, or Debian-style ifupdown and binds the profile to the preserved MAC. On EL7, whose NetworkManager ignores keyfile profiles unless configured to load them, it uses a MAC-bound ifcfg profile. With ifupdown (networking.service and /etc/network/interfaces), which cannot match a MAC, it replaces the overlay's /etc/network/interfaces with a static stanza for the guest's configured interface and keeps the original beside it as interfaces.certistack-original. Bridges, bonds, VLANs and tunnels are set aside in favor of the physical interface under them. When the guest names several physical interfaces and the sandbox VM has only the one NIC (for example cloud-init's enp6s18 beside an image's stale eth0), each name gets an allow-hotplug stanza and only the one that exists comes up; with several NICs (attach_all_nics) it asks for interface_name. It assigns only the isolated address; it creates no gateway, default route, DNS setting, SNAT, or physical uplink. The source VM, PBS snapshot, and read-only mapped image are never changed. The signed VM record lists the chosen strategy, address, and generated guest path.

When a guest has firewalld installed, the same disposable overlay receives a copy of that guest's effective default zone with source-restricted rules added: the sandbox probe address may reach the configured Stage 2 probe ports and protocols. Everything the guest's own zone already allowed (services, ports, rich rules, target) is kept. This makes an otherwise-correct restored service testable even when its own firewall would refuse the probe, without opening a production firewall or mutating the backup. The change disappears with the COW overlay during cleanup.

Restored guests on the same isolated VNet therefore reach each other exactly as far as their own firewalls allow, as they do in production. A tier that reports on its dependencies (for example a readiness endpoint that checks its peers) can be asserted by a worker probe when the tiers are restored in dependency order. CertiStack adds no rule for guest-to-guest traffic. Routed multi-VLAN topology and application-transaction evidence remain outside the supported boundary until a separate, explicitly qualified isolated-network capability exists.

Set capture_diagnostics: true for a compatibility investigation. On a wire probe failure, CertiStack records the best safe guest-side evidence available (interfaces, routes, listening sockets, and relevant service state) in the signed report. Guest-agent diagnostics are optional: an unavailable or restricted guest agent is reported as such and does not turn a passing wire probe into a failure.

interface_name is optional. Use it only for a confirmed legacy profile that is tied to an interface name (for example ifcfg-eth0); CertiStack adds a MAC-specific .link file alongside the temporary recovery profile so the guest sees that name in the sandbox. It never renames a source interface.

Only net0 is attached by default. Set attach_all_nics: true to attach the source's other NICs (net1 and up) as well: each keeps its MAC and model and is attached, like net0, only to the isolated VNet with the PVE firewall on. Their bridges, VLAN tags and firewall settings are never replayed. Use it when a service binds to a secondary interface's address and would not start without it. Only net0 gets a recovery address; the other NICs keep the guest's own configuration, and a guest that waits for DHCP on them boots more slowly, because the sandbox has no DHCP. Those NICs bring up their production static addresses on the sandbox VNet, so use the option only when none of them falls inside network.ip_range: an address there could answer another VM's probe or clash with probe_ip. The report lists attached NICs under restored_hardware.additional_nics and the ones left out under omitted_nics.

To probe a service on another NIC's address, set sandbox_ip to net0's address and point the probe at the other NIC's address. That address is the guest's own, so it must be inside network.ip_range to be reachable. It must also be unique in the plan: validation refuses an address another VM uses. Whether the guest's firewall admits the probe on that interface depends on the guest's configuration of it.

network_recovery:
  sandbox_ip: "10.0.70.75"       # net0, configured by network recovery
  attach_all_nics: true
probes:
  - {type: http, target_ip: "10.0.70.75", port: 80, path: /healthz, expected_status: 200, timeout_sec: 120, description: "nginx on net0"}
  - {type: tcp, target_ip: "10.0.71.202", port: 8080, timeout_sec: 120, description: "service bound to net1's address"}

Omitted network_recovery means auto; CertiStack infers the one guest address from that VM's Stage 2 probe targets and detects the supported Linux network backend. Use preserve explicitly to opt out. A plan with wire probes for more than one guest address fails before recovery rather than silently guessing which guest identity to rewrite.

Windows and other non-Linux guests

Guest network recovery rewrites Linux network configuration on the disposable overlay. It cannot adapt a Windows guest, or any guest whose network configuration it does not recognise, and a VM with wire probes and the default mode then fails with an unsupported guest network layout error that names network_recovery.mode: preserve. For such a guest:

  • set network_recovery.mode: preserve, so the guest keeps its own address on the isolated VNet;
  • make network.ip_range the subnet the guest already uses, and each wire probe's target_ip that guest's own address (a disconnected L2 domain can reuse production addressing; see the isolation attestation in the supported recovery contract);
  • do not set sandbox_ip, which preserve forbids.

Guest-agent probes (qga with sc query <service>) need no guest network at all. The Windows domain example puts both to work.

use_backup_config: true remains the recommended path: it discovers all backed-up disks, safe hardware settings, and the preserved net0 identity. When full inheritance is intentionally disabled, a run still reads what a restore cannot do without (firmware, boot order, machine type, SCSI controller, guest OS type, the backed-up disk list; see "Source disk layout"), and an automatic wire-probe run also reads the net0 MAC/NIC identity for the disposable profile. It does not import compute settings, source network topology, or source tags.

The temporary node worker needs guestfish from libguestfs-tools to make this overlay-only change. The appliance image/Enterprise connector should include that dependency. Community reports a clear preflight-style error if it is unavailable; it never falls back to changing a source guest or attaching it to production networking.

The signed report separates this offline COW guest-network preparation from mapping and VM boot time, both per VM and in the RTO timing breakdown. This makes guest-layout compatibility costs visible without weakening isolation. "Boot duration" begins at the sandbox VM start request and ends at Stage 1 hypervisor readiness; it does not include storage mapping, guestfs work, or full mapped-image integrity verification.

For a passing validation, rto_seconds and phase_timings.total_rto_sec measure time from the run start through the last verified recovery probe; teardown is excluded and reported separately as teardown_sec. A failed validation has no achieved recovery point, so those RTO fields measure elapsed time through terminal failure, including teardown when cleanup runs before the result is finalized. Compare failed-run cleanup duration using teardown_sec, not by treating its RTO as a successful recovery objective.


3. Deterministic Probing Specifications

Type: tcp (Stage 2: Wire-Level TCP Handshake)

Executes an active TCP 3-way handshake from the node worker's temporary VNet address against the guest IP and port. Stage 2 TCP/HTTP/DNS/LDAP evidence proves worker-to-guest reachability, not an in-guest client path; each signed result records its execution origin. Use a QGA probe when the assertion must execute inside the guest.

- type: "tcp"
  target_ip: "192.0.2.10"
  port: 5432
  timeout_sec: 15
  description: "Verify PostgreSQL socket listener on port 5432"

Type: http (Stage 2: Wire-Level HTTP/HTTPS Health Endpoint)

Transmits an HTTP GET request, verifies TLS handshake, and asserts response status code.

- type: "http"
  target_ip: "192.0.2.30"
  port: 443
  path: "/healthz"
  tls: true
  server_name: "service.example.com" # optional: TLS SNI and HTTP Host header
  expected_status: 200
  timeout_sec: 20
  description: "Verify Web App HTTPS healthz endpoint returns 200 OK"

HTTPS probes always verify the certificate chain and hostname. The recovered service certificate must chain to a CA trusted by the node worker and its SAN must include server_name (or the target IP when no server name is supplied). insecure_skip_verify and tls_skip_verify are rejected deliberately: a successful recovery assertion must not be based on an unauthenticated TLS endpoint. For an isolated lab, install the lab CA on the node worker rather than disabling verification.

Type: dns (Stage 2: Wire-Level DNS Resolution)

Sends one DNS query directly to target_ip over UDP (re-sent until the timeout) or TCP. The worker's own /etc/hosts and resolver configuration are never consulted.

  • With domain, the probe passes only when the server answers NOERROR with at least one A or AAAA record for that name. NXDOMAIN, REFUSED, SERVFAIL, or an empty answer fails, so a restored DNS server that lost its zone is detected.
  • Without domain, it is a listener check: the server must return any well-formed reply to a root NS query.
- type: "dns"
  target_ip: "192.0.2.10"
  port: 53
  domain: "healthy.example.com"
  transport: "tcp" # optional; defaults to udp
  timeout_sec: 10
  description: "Verify internal DNS listener responds to queries"

With record_type: srv, the probe asks for SRV records instead and passes only when the answer holds at least one. The report lists each target, port, priority and weight. For a domain controller, the DC locator record _ldap._tcp.dc._msdcs.<domain> proves the restored DNS server loaded the AD-integrated zone, which it reads from the directory, and still carries the records clients use to find a DC. The records were written before the backup, so they do not show that Netlogon re-registered after the restore.

- type: "dns"
  target_ip: "192.0.2.10"
  port: 53
  domain: "_ldap._tcp.dc._msdcs.example.com"
  record_type: "srv"
  timeout_sec: 60
  description: "Verify the DC locator records are served"

Type: ldap (Stage 2: Wire-Level Active Directory / Directory Probe)

Performs an LDAP v3 anonymous bind over TCP port 389. When base_dn is set, it also performs an anonymous base-object search for that DN.

Use an explicit DN when it is a business assertion that must exist. CertiStack does not silently weaken an explicit DN if the restored directory advertises a different naming context. For a portable service-recovery check across heterogeneous customer directories, set base_dn: auto: CertiStack reads the server's Root DSE and requires a successful base-object search of one of its advertised naming contexts. Omit base_dn only when an LDAP v3 listener/bind handshake alone is the intended assertion.

- type: "ldap"
  target_ip: "192.0.2.10"
  port: 389
  base_dn: "dc=example,dc=com" # explicit assertion; requires anonymous search access
  timeout_sec: 15
  description: "Verify Active Directory domain controller LDAP listener"
- type: "ldap"
  target_ip: "192.0.2.10"
  port: 389
  base_dn: "auto" # portable Root-DSE discovery plus verified base-object search
  timeout_sec: 15
  description: "Verify any restored directory naming context is readable"

Active Directory. A domain controller refuses anonymous searches by default, so base_dn (explicit or auto) fails against it, and omitting base_dn only proves that something answers a bind. Set naming_context instead: CertiStack binds anonymously, reads the Root DSE (the one entry AD lets anonymous clients read), and passes only if the directory advertises that naming context, as defaultNamingContext or in namingContexts. No directory search is sent and the DC's policy stays as it is. The directory service answers the Root DSE itself, so a DC whose database did not open fails even though port 389 accepts connections. The comparison ignores case and the spaces around , and =. The report records the DC's dnsHostName and isSynchronized. isSynchronized is informational: a DC restored without its replication partners may report FALSE while serving its directory correctly. naming_context works for OpenLDAP and 389-ds too (their Root DSE lists namingContexts), and cannot be combined with base_dn.

- type: "ldap"
  target_ip: "192.0.2.10"
  port: 389
  naming_context: "DC=example,DC=com" # Root DSE assertion; no directory search
  timeout_sec: 60
  description: "Verify the restored domain controller serves example.com"

Type: mssql (Stage 2: Wire-Level SQL Server Protocol Probe)

Sends a TDS PRELOGIN, the first message of every SQL Server connection, and passes only on a well-formed PRELOGIN response. No login is attempted and no credentials are needed, so the probe proves that the SQL Server engine answers its protocol, not merely that the port is open. The report records the server's version and encryption setting. A named instance must listen on a static port, because the probe does not query SQL Browser. A server configured for TDS 8.0 strict encryption expects TLS before any TDS message, and the probe reports that instead of passing.

- type: "mssql"
  target_ip: "192.0.2.20"
  port: 1433
  timeout_sec: 120
  description: "Verify SQL Server answers TDS on the restored guest"

Type: smb (Stage 2: Wire-Level SMB Protocol Probe)

Sends an SMB2 NEGOTIATE over direct TCP and passes only on a successful NEGOTIATE response for one of the offered dialects (SMB 2.0.2 through 3.0.2). The report records the dialect and whether the server requires signing. No session is set up and no share is opened, so the probe proves the SMB server answers, not that a particular share or its data survived. For that, pair it with an application check that reads the data.

- type: "smb"
  target_ip: "192.0.2.21"
  port: 445
  timeout_sec: 60
  description: "Verify the file server answers SMB on the restored guest"

Type: qga (Stage 3: In-Guest Inspection via QEMU Guest Agent)

Uses a constrained QGA health-check vocabulary inside the guest operating system. Plan-authored commands are limited to hostname, guest-get-host-name, ping, guest-ping, systemctl is-active <service>, and sc query <service>; arbitrary shell commands and environment interpolation are rejected before execution. Guest output lands in the signed report, so the vocabulary stays read-only and narrow.

- type: "qga"
  command: "systemctl is-active postgresql"
  expected_output: "active"
  require_qga: true # make a missing/unresponsive guest agent a failure
  timeout_sec: 30
  description: "Verify PostgreSQL service unit state is active inside guest"

On Windows guests, sc query <service> checks a service by its key name (for example W3SVC, LanmanServer, or MSSQL$SQLEXPRESS). CertiStack runs sc.exe directly with the literal arguments query and the name, never through a shell, and passes only when the service's STATE is 4 RUNNING. A stopped, pending, or missing service fails, and the report quotes the state or sc.exe's error. expected_output is optional and adds a substring check. The name may contain letters, digits, and $ _ . -. The Windows guest agent must allow guest-exec, which it does by default.

- type: "qga"
  command: "sc query MSSQL$SQLEXPRESS"
  require_qga: true
  timeout_sec: 60
  description: "Verify SQL Server Express is running inside the restored guest"

Fallback Invariant: If QGA is not installed or unresponsive in the guest, Stage 3 probes are marked as SKIPPED and pass/fail criteria falls back strictly to Stage 2 wire verification.

Set require_qga: true for a test whose assertion is the guest-agent channel itself. In that case an unavailable QGA channel is a probe failure instead of a skipped optional inspection.

Type: qmp (Stage 1: Hypervisor State)

Connects to the sandbox VM's QMP socket and passes only when query-status reports the VM running. It proves the VM is powered on, not that its guest operating system or any service works; use it for a VM that has nothing else to probe, such as one without a guest agent. The alias qmp_status is accepted. The result records its origin as hypervisor_qmp.

- type: "qmp"
  timeout_sec: 60
  description: "Verify the sandbox VM is running"

Type: screendump

Captures the VM's display through QMP and passes when the capture succeeds. It asserts nothing about what the screen shows; it is for the record, in the same way as the forensic capture CertiStack takes when any probe fails.

- type: "screendump"
  timeout_sec: 30
  description: "Capture the console after boot"

4. Literal Plans and Protected Runtime Environment

Plans are literal documents. CertiStack does not expand ${VAR_NAME} or shell syntax in a plan: a plan can be submitted through an API, retained as signed evidence, and inspected by another operator, so process-environment interpolation would risk credential disclosure and make its effective content ambiguous.

Keep connection credentials in a protected allowlisted environment file passed with --env-file; its values are parsed literally, never shell-sourced. Keep plan identity and controller-local signing-key paths explicit:

plan_id: "acme-erp-dr"
compliance:
  sign_key_path: "/secure/certistack/controller-signing.ed25519"
vms:
  - vmid: 9001
    source_vmid: 100
    source_pbs:
      storage_id: "pbs"
      snapshot: "latest"

If a site needs a family of plans, render them in a controlled build step, validate each generated file, and retain its SHA-256 in the campaign's plans.tsv. Do not run a shell template against a secret-bearing environment.


5. Complete field reference

The parser rejects unknown keys, a plan larger than 10 MiB, and a file that contains more than one YAML document, so every key it accepts is listed here. An unknown key's error names its line and section, the key you probably meant and the keys that section accepts, for example line 4: unknown key "grace" in tiers[]; did you mean "startup_grace_period_sec"?. An alias is accepted for older plans and normalised to the canonical key; write the canonical key in new plans. A rejected key parses but fails validation, so an unimplemented or unsafe setting can never be silently ignored. Credentials never belong in a plan: pbs_password, tokens, and similar keys are unknown to the parser and fail to load.

Top level

Key Type Status Rule
version string required v1; certistack/v1 is an accepted alias
plan_id string required matches ^[a-zA-Z0-9_-]+$
name string required non-empty
environment string optional free-text label recorded in the report and shown by validate
use_backup_config bool optional inherit safe compute settings and all backed-up disks from each resolved snapshot
compliance map required see below
admission_control map required see below; admission is an accepted alias
network map required see below
storage map optional, discouraged see below
tiers list required at least one tier

compliance

Key Type Status Rule
frameworks list of string optional may be empty; each item one of DORA, DORA-Article-12, SOC2, SOC2-Type2-CC7.4, HIPAA, HIPAA-164.308-A7, ISO27001, ISO27001-A8.13, ISO27001-A.12.3, ISO27001-A12.3, CIS_V8
retention_days int optional 0 (the default) to 36500; recorded in the report, never enforced
sign_key_path string required on the controller controller-local Ed25519 private key; never sent to the worker; --key overrides it
rto_target_seconds float optional >= 0; when set, a run whose measured RTO exceeds it is reported as failed

admission_control

Key Type Status Rule
max_host_ram_percent float required > 0 and <= 100
max_host_iowait_percent float required > 0 and <= 100; max_host_io_wait_percent is an accepted alias
defer_retry_interval_sec int required > 0
max_defer_retries int required > 0
max_concurrent_boots int optional 0 (unlimited, the default) to 64
boot_settle_sec int optional 0 (default 180) to 3600; only with max_concurrent_boots

network

Key Type Status Rule
mode string required simple or vxlan
lifecycle string optional ephemeral (default) or preprovisioned
zone_id, vnet_id string required PVE SDN identifier: a letter followed by letters and digits
tag int required for vxlan VNI 1..16777215, unused by every existing PVE VNet; optional for simple
peers string required for vxlan comma-separated IP addresses without duplicates; rejected for simple
ip_range string required IPv4 CIDR of the isolated subnet, outside loopback, link-local, multicast, and reserved space; subnet_cidr is an accepted alias
probe_ip string required address inside ip_range that is neither the network nor the broadcast address
gateway string rejected must be empty; a sandbox never has a gateway
ipam string rejected not implemented; must be empty
mtu int rejected not implemented; must be unset

storage

Prefer the protected environment file (PBS_REPOSITORY, PBS_FINGERPRINT) over this stanza: a repository string can contain a token identity and a plan is retained as signed evidence.

Key Type Status Rule
pbs_repository string optional, discouraged when set, a VM may omit source_pbs and defaults to vm/<vmid>/latest
pbs_fingerprint string optional, discouraged PBS certificate fingerprint
pbs_namespace string optional default for every VM's source_pbs.namespace
scratch_dir string optional node-local private work directory (guest network recovery, screendumps; also overlays when overlay_storage is unset); the worker rejects shared temporary directories such as /tmp
overlay_storage string required on stock Proxmox VE PVE storage ID of a file-based storage (dir, nfs, cifs, cephfs, btrfs) with images content. Overlays are created as <storage>/images/<vmid>/certistack-delta-*.qcow2 and attached by volume ID. The run fails before any mutation if the storage is inactive, lacks images, or has under 5 GiB free, and it refuses an images/<vmid> directory that holds other volumes. --overlay-storage / CERTISTACK_OVERLAY_STORAGE override it
max_copy_gib int optional 0 (no cap, the default) to 1048576. The most copy-before-boot may write in one run, all VMs together. auto does not copy a VM whose copies would pass it, and warns; always fails that VM. Use it where the free space is not all yours, such as a scratch disk shared with other tenants

tiers[]

Key Type Status Rule
level int required >= 1, unique across the plan, executed in ascending order; tier is an accepted alias
name string optional defaults to Tier <level>
startup_grace_period_sec int optional >= 0
boot_timeout_sec int optional >= 0; default 60; PVE start plus QMP readiness
soak_time_sec int optional >= 0; hold the tier for this long after its probes pass before starting the next tier; recorded in the report timings
vms list required at least one VM

tiers[].vms[]

Key Type Status Rule
vmid int required > 0, unique across the plan; the temporary sandbox VM; target_vmid is an accepted alias
source_vmid int required when snapshot: latest the protected source VM
name string required non-empty
mac string optional six-octet MAC for the isolated NIC; normally inherited with use_backup_config
nic_model string optional virtio, e1000, e1000e, igb, ne2k_pci, ne2k_isa, pcnet, rtl8139, vmxnet3
machine, ostype, bios, scsihw string optional explicit PVE values that override inherited ones
network_recovery map optional see below
source_pbs map required unless storage.pbs_repository is set see below
hardware_overrides map optional cores, sockets, memory_mb, balloon_mb, cpu_weight (ints)
copy_before_boot string optional auto (default), always or never. Whether the integrity scan also copies the VM's disks next to its overlays, so the sandbox boots from local copies as a restore does. auto copies when the tier's disks would not all stay cached until the VM boots, unless overlay storage is on the cluster database's disk or lacks room for the disks plus 5 GiB; it warns when it can't. always copies or fails the VM. Reports record copied_before_boot and each disk's boot_source
require_quiesced_backup bool optional default false; fail the VM before anything is restored unless its backup is quiesced or powered-off (see Consistency above)
probes list required at least one probe

tiers[].vms[].source_pbs

Key Type Status Rule
storage_id string required pbs (this run's configured endpoint); other selectors are rejected at run time
snapshot string required latest, vm/<vmid>/latest, or vm/<vmid>/<RFC 3339 time>
drives list required unless use_backup_config entries of slot (scsi0..scsi30, virtio0..virtio15, sata0..sata5, ide0..ide3; unique) and archive (PBS archive filename without path separators); every slot must be backed up in the snapshot
drive string deprecated alias a single archive, in the slot it names (drive-virtio0) or scsi0
namespace string optional PBS namespace; defaults to storage.pbs_namespace

tiers[].vms[].network_recovery

Key Type Status Rule
mode string optional auto (default), preserve, ifcfg, networkmanager, netplan, systemd-networkd, ifupdown
sandbox_ip string optional usable IPv4 address in network.ip_range, unique per plan; required when a VM's wire probes target more than one address; forbidden with preserve
interface_name string optional Linux interface name (not lo) for a legacy name-bound profile
capture_diagnostics bool optional record guest-side diagnostics on a wire-probe failure
attach_all_nics bool optional also attach the source's net1+ NICs (MAC and model only) to the isolated VNet

tiers[].vms[].probes[]

Key Type Status Rule
type string required tcp, http, dns, ldap, mssql, smb, qga, qmp, screendump; aliases tcp_handshake, http_get, qga_ping, qmp_status
timeout_sec int required > 0 (a missing value becomes 10)
retries int optional 0..10; 0 keeps automatic polling until the timeout
description string optional defaults to <type> verification probe
target_ip string required for tcp, http, dns, ldap, mssql, smb literal IP inside network.ip_range, never probe_ip or the subnet network or broadcast address; a run refuses a target that is an address of the worker host (for example another host address on a preprovisioned VNet)
url string optional (http) http(s)://host[:port]/path; expanded into target_ip, port, path, and tls when target_ip is empty
port int required for tcp, dns, ldap, mssql, smb > 0
path, tls, server_name, expected_status mixed http expected_status must be > 0; server_name must be a valid DNS name and sets both SNI and the Host header
tls_skip_verify, insecure_skip_verify bool rejected any true value fails validation
domain, transport string dns transport is udp (default) or tcp
record_type string optional (dns) a (default: an A or AAAA answer) or srv; srv needs domain
base_dn string optional (ldap) explicit DN, auto, or omitted
naming_context string optional (ldap) a DN the Root DSE must advertise; not with base_dn
command, expected_output, require_qga mixed qga command limited to hostname, guest-get-host-name, ping, guest-ping, systemctl is-active <service>, or sc query <service> (Windows)