Declarative Test Plan Specification & Reference
CertiStack test plans are written in YAML (version: "v1"). Sections 1–4 describe the recommended plan layout, probe definitions, and the literal-plan rule. Section 5 lists every key the parser accepts, treats as a deprecated alias, or deliberately rejects.
1. Top-Level Schema Structure
version: "v1"
plan_id: "string (alphanumeric, hyphens, underscores)"
name: "Human readable plan title"
# Optional: inherit safe compute settings and all backed-up disks from
# each resolved PBS snapshot. Network and PVE ownership settings stay isolated.
use_backup_config: true
compliance:
frameworks: [ "DORA-Article-12", "SOC2-Type2-CC7.4", "HIPAA-164.308-A7", "ISO27001-A8.13" ] # optional
retention_days: int (0 to 36500; optional)
sign_key_path: "string (/path/to/private.key)"
admission_control:
max_host_ram_percent: float (> 0 and <= 100)
max_host_iowait_percent: float (> 0 and <= 100)
defer_retry_interval_sec: int (> 0)
max_defer_retries: int (> 0)
network:
mode: "simple" | "vxlan"
zone_id: "string"
vnet_id: "string"
tag: int (VXLAN VNI 1..16777215; required for vxlan)
peers: "comma-separated peer IPs (required for vxlan)"
ip_range: "CIDR string (e.g. 192.0.2.0/24)"
probe_ip: "host address inside ip_range (e.g. 192.0.2.1)"
tiers:
- level: int (>= 1, executed in ascending order)
name: "string"
startup_grace_period_sec: int
boot_timeout_sec: int (>= 0; default 60; applies to PVE start and QMP readiness)
vms:
- vmid: int (> 0, unique across entire plan; temporary test VM)
source_vmid: int (> 0; protected source VM)
name: "string"
source_pbs:
# `pbs` selects this run's configured PBS endpoint. Other selectors
# are rejected instead of silently routing a VM to a different backup.
storage_id: "pbs"
snapshot: "latest" | "pinned PBS snapshot path"
drives:
- slot: "scsi0"
archive: "drive-scsi0"
hardware_overrides:
cores: int
sockets: int
memory_mb: int
balloon_mb: int
cpu_weight: int
probes:
- type: "tcp" | "http" | "dns" | "ldap" | "mssql" | "smb" | "qga"
timeout_sec: int (> 0; total probe readiness budget)
# Optional. 0 keeps automatic polling until timeout; 1-10 limits
# the probe to that many additional executions.
retries: int (0 to 10)
...
When use_backup_config: true, CertiStack discovers the backed-up disk layout
and safe compute/firmware settings from qemu-server.conf. It also preserves
the source net0 MAC and NIC model while attaching that NIC only to the
isolated VNet; this keeps MAC- and driver-bound guest network profiles working
without replaying a production bridge, firewall setting, VLAN, gateway, or
source PVE tags.
2. Configuration Field Definitions
compliance
Controls audit metadata and non-repudiation signing.
| Field | Type | Required | Description |
|---|---|---|---|
frameworks |
[]string |
No | Target regulatory standards recorded in the report; the list may be empty. Canonical items: DORA-Article-12, SOC2-Type2-CC7.4, HIPAA-164.308-A7, ISO27001-A8.13; section 5 lists the short forms and variants that are also accepted. A framework label is not a certification. |
retention_days |
int |
No | Audit record retention the report states (e.g., 2555 for 7-year financial retention), from 0 (the default) to 36500. It is recorded in the signed report for whoever keeps it; CertiStack never prunes or deletes a report. |
sign_key_path |
string |
Yes | Controller-local path to the Ed25519 private key used to sign the final audit certificate. The path is never sent to the PVE node worker. --key or CERTISTACK_SIGN_KEY overrides it at run time, but validate still requires it in the plan. |
admission_control
Protects production workloads running on the hypervisor host.
The first four keys are required: a plan that omits one fails validation. The
values in the examples are the ones shown below. max_concurrent_boots and
boot_settle_sec are optional and have the defaults shown.
| Field | Type | Required / default | Description |
|---|---|---|---|
max_host_ram_percent |
float |
required; examples use 85.0 |
Maximum host RAM utilization allowed before starting a test. |
max_host_iowait_percent |
float |
required; examples use 12.0 |
Maximum host CPU I/O wait % allowed before starting a test. |
defer_retry_interval_sec |
int |
required; examples use 30 |
Seconds to sleep between admission re-evaluations. |
max_defer_retries |
int |
required; examples use 10 |
Maximum retry attempts before marking the test run as throttled/failed. |
max_concurrent_boots |
int |
optional, 0 (unlimited) |
How many sandbox VMs may boot at once on the node. Before starting a VM, the run waits while this many VMs started within boot_settle_sec. |
boot_settle_sec |
int |
optional, 180 |
How long a started sandbox VM counts as booting for max_concurrent_boots. |
A tier starts its VMs one after another, as soon as each is prepared, so
their first minutes of boot overlap. That is when each guest writes most to
its new overlays. On a node whose overlay storage shares a disk with the
cluster file system's database (the run warns about it), several overlapping
boots can delay corosync and cost the node its cluster membership.
max_concurrent_boots: 1 or 2 spaces them out, across tiers too. The wait
is logged, counted in the run's admission time, and shown by certistack
plan before each start it may hold.
In release mode, each VM's pre-boot integrity scan reads all of its disks before it starts, so the starts already come about one scan apart: in the lab, a tier of three VMs with 10-minute scans never waited. The limit holds starts back mostly for VMs with small disks, whose scans are short, and in test mode, which has no scan.
Before creating any sandbox resource, CertiStack also reserves the declared maximum memory for the largest concurrent recovery tier plus 256 MiB per sandbox VM for host overhead. It combines that reservation with current host usage and the RAM percentage cap. Ballooning is not counted as guaranteed free capacity. This fails safely before PBS mapping when a full or OS-wide matrix would overcommit the node.
network
Configures isolated Proxmox Software-Defined Network (SDN) sandboxes.
| Field | Type | Options | Description |
|---|---|---|---|
mode |
string |
simple, vxlan |
simple for single-node air-gapped bridges; vxlan for multi-node clusters. |
lifecycle |
string |
ephemeral (default), preprovisioned |
ephemeral makes CertiStack create and remove the zone/VNet. preprovisioned uses a platform-owned sandbox only after read-only validation of its zone, VNet relationship, subnet, no-gateway, and no-SNAT invariants. |
zone_id |
string |
— | Unique PVE SDN identifier (e.g. csZone01; letters and digits only, starting with a letter). |
vnet_id |
string |
— | Unique PVE SDN identifier (e.g. csVnet01; letters and digits only, starting with a letter). |
tag |
int |
1..16777215 for vxlan |
VXLAN VNI assigned to the VNet. It must be explicit and unused by every existing PVE VNet; simple networks may omit it. |
peers |
string |
comma-separated IP addresses for vxlan |
Required VXLAN zone peer list. CertiStack canonicalizes the addresses and sends them as the PVE zone peers field; simple networks must omit it. |
ip_range |
string |
CIDR | Isolated subnet CIDR block (e.g. 192.0.2.0/24). Strictly no gateway or SNAT is provisioned. It must not overlap loopback, unspecified, link-local, multicast, or reserved space, which the worker host would answer itself. |
probe_ip |
string |
IP address in ip_range |
Required temporary address assigned by the node worker to the PVE VNet bridge for Stage 2 probes. It must not be the subnet network/broadcast address or a guest target. The worker requires per-interface IPv4 forwarding to be disabled and installs scoped host-input/forwarding containment before assignment. The address adds only the connected subnet route and is removed during teardown/recovery; it is never a gateway or SNAT address. Wire probes (tcp, http, dns, ldap, mssql, smb) are sent from this address and bound to the VNet interface, so neither the host routing table nor a host listener can answer them. |
The IDs may remain stable across scheduled runs. With the default
lifecycle: ephemeral, a successful CertiStack run removes the resources it
owns. Before it creates anything, CertiStack reads the exact zone and VNet IDs.
If either already exists, it fails closed without reusing, modifying, or
deleting it. This prevents a validation plan from adopting a customer or a
previously interrupted sandbox. Review the matching PVE resource or reconcile
the interrupted run, then retry the same plan; do not work around the condition
by repointing a production configuration.
Applying SDN in PVE is all or nothing, on every node, so CertiStack creates
and removes its zone and VNet under PVE's SDN configuration lock and applies
only its own change. PVE refuses the lock while the cluster's SDN has changes
nobody applied, and the run is then refused before it creates anything
(troubleshooting Issue 3d). plan reports it too. The lock is held for
seconds, around each create-and-apply and remove-and-apply, and needs
SDN.Allocate on /sdn, which the documented role grants.
Use lifecycle: preprovisioned when the platform operator deliberately provisions one
dedicated, air-gapped sandbox ahead of time. CertiStack performs only read-only
SDN API requests for that sandbox and retains it at teardown; it still adds and
removes the exact host-side probe address and destroys every test VM, overlay,
and PBS map it owns. A report records the validated-and-retained sandbox rather
than falsely claiming it was purged.
PVE's ifupdown2 turns IPv4 forwarding on for every SDN bridge at each
network reload. CertiStack refuses a sandbox VNet that forwards, because
that would let the host route between the sandbox and other networks. For
an ephemeral sandbox, the bridge is the run's own: the run waits for the
node's network reload to bring it up and turns forwarding off itself. The
containment table it adds with the probe address drops forwarded packets as
well, even if a later network reload turns forwarding back on. A
pre-provisioned VNet is the platform's, so the platform keeps forwarding off
on it, for example with an /etc/network/if-up.d/ hook that writes 0 to
/proc/sys/net/ipv4/conf/<vnet>/forwarding for that VNet on every node that
may run a validation. A run otherwise refuses it before adding the probe
address.
tiers & Dependency DAG
Tiers define the sequential boot order of your application architecture. Tiers are sorted by level in ascending order ($1 \to 2 \to 3$).
- CertiStack starts each planned VM in Tier $N$, waits for
startup_grace_period_sec, and evaluates Stage 2 & Stage 3 probes for every VM that reached Stage 1. - A failure does not hide its independent peers: CertiStack collects every current-tier outcome and a failure framebuffer per affected VM in the signed report.
- Only after all VMs and probes in Tier $N$ pass does Tier $N+1$ begin execution. A failed tier blocks dependent tiers.
Snapshot selection
Set each VM's source_vmid and source_pbs.snapshot: "latest" for normal
scheduled validation. Before any SDN or VM resource is created, CertiStack
queries that VM's PBS backup group, selects the highest backup-time, and
replaces latest with the exact immutable snapshot path. That resolved path is
included in the report. A full snapshot path remains supported only when an
operator intentionally needs to replay a historical recovery test.
Source disk layout
Use source_pbs.drives for every backed-up disk. Each entry keeps its source
PVE slot on any bus (scsi0–scsi30, virtio0–virtio15, sata0–sata5,
ide0–ide3) and names the PBS archive. Do not include Cloud-Init CDs or
other non-backed-up media. source_pbs.drive remains accepted for existing
single-disk plans; its slot is the one its archive names (drive-virtio0 is
virtio0), else scsi0.
Every run reads the snapshot's qemu-server.conf, whatever the plan
inherits, and applies what a restore cannot do without:
- Boot-critical hardware. Firmware (see below), machine type, SCSI controller and guest OS type come from the backup unless the plan sets them, because a guest may not boot on PVE's defaults: a Windows guest installed on an LSI controller has no driver for virtio-scsi.
- Machine version as a restore resolves it. Proxmox VE pins a new
Windows VM's machine version on create (
q35becomes, for example,pc-q35-11.0+pve2), which a restore does not:qmrestorekeeps the backup's setting, and PVE resolves an unversioned one at every start. The sandbox VM gets the backup's setting back after create, and the report records the version it actually ran asrestored_hardware.running_machine. For a Windows guest, PVE starts an unversioned machine type on the version of the QEMU that created the VM (itsmetaline), or on 5.1 when that QEMU is older than 9.1. PVE writes the sandbox VM'smetaitself and refuses to change it, so when the backup's version differs from the node's latest, the sandbox VM is pinned explicitly to the version a restore would start on. - SMBIOS identity as a restore keeps it. The sandbox VM gets the backup's
smbios1(its UUID and any manufacturer, product, serial, SKU and family strings), asqmrestorekeeps it; a guest can tell a new SMBIOS UUID from its own. The VM generation ID is new, as after any restore. The report records the UUID asrestored_hardware.smbios_uuid. - No missing drives. A restore keeps the source's CD-ROM drives (a
cloud-init drive, an ISO), so the guest finds the same controllers. The
sandbox VM gets them empty, so no configuration drive is regenerated with
the source's network settings. The report lists them as
restored_hardware.empty_cdroms. - No new devices. A source without a memory balloon device
(
balloon: 0) gets none in the sandbox either, unless the plan setshardware_overrides.balloon_mb; a device the guest never had is new hardware to it. The report records this asrestored_hardware.no_balloon. - Boot disk. The sandbox VM boots the first restored disk in the source's
boot order (
boot: order=..., or the legacybootdisk). Without one it boots the disk PVE would pick: the first by bus (ide,scsi,virtio,sata) and index. The report records it asrestored_hardware.boot_disk. - Disks the backup has. A plan that lists a disk the backup does not contain is refused.
- Disks the plan leaves out. A plan may restore only some of the
backed-up disks, for example the boot disk alone. The disks it leaves out
are listed in the VM's
omitted_source_disksin the signed report, and the run warns, so a partial restore is never presented as the whole VM. Each archive must go to the slot it was backed up from, and only once. - Disks no backup contains. Disks the source excludes from backups
(
backup=0) are listed asexcluded_source_disks: no restore can include them, so the evidence says the VM was never whole in PBS. - Protection.
source_crypt_modeis the weakest PBS crypt mode among the VM's disk archives (encrypt,sign-onlyornone), andsource_signature_verifiedis true when the configured key verified the snapshot manifest's signature. - Consistency.
source_consistencysays how consistent the backed-up disks were, from the backup's own log (client.log.blob, which vzdump stores with the snapshot): quiesced: the guest agent froze the file systems (VSS on Windows);crash-consistent: a running VM was backed up without a freeze that held (the agent was missing or not running, the freeze was refused, for example by SELinux, or Windows released it before the snapshot);powered-off: the VM was not running, or was shut down for the backup;unknown: the log could not be read.source_consistency_detailgives the reason in the log's words, and a crash-consistent backup is flagged in the run's progress. It fails a VM only when the plan setsrequire_quiesced_backup: trueon it: then anything butquiescedorpowered-off,unknownincluded, fails that VM before it is restored, anddoctorfails it too. The log is read through the PBS API with the configured repository, credentials and certificate fingerprint, becauseproxmox-backup-clientcannot fetch a file outside the snapshot manifest. It is not covered by the manifest's signature, and an encrypted backup's log is reported asunknown.
source_pbs:
storage_id: "pbs"
snapshot: "latest"
drives:
- slot: "scsi0"
archive: "drive-scsi0"
- slot: "scsi1"
archive: "drive-scsi1"
Backup configuration inheritance
Set use_backup_config: true to have CertiStack read the exact resolved
snapshot's qemu-server.conf and discover every disk image archive
(drive-scsiN, drive-virtioN, drive-sataN, drive-ideN). The inherited
allowlist is CPU cores/sockets, memory/ballooning, CPU weight, machine type,
guest OS type, firmware, and SCSI controller. Explicit plan
values override these individual inherited fields.
CertiStack never replays source network interfaces, PVE tags, hooks, startup rules, host PCI/USB passthrough, or arbitrary QEMU arguments. Source tags are preserved as attestation evidence only. It creates a fresh isolated NIC and applies only CertiStack ownership tags.
UEFI and TPM firmware state
A UEFI source VM (bios: ovmf) keeps its firmware state in two small disks:
the UEFI variable store (efidisk0: boot entries and Secure Boot keys) and,
with a virtual TPM, the TPM state (tpmstate0). CertiStack restores both
with every plan, whether or not it sets use_backup_config, because a UEFI
VM restored without them, or under SeaBIOS, is a different machine:
- Each image is restored whole from PBS into a new raw volume on
storage.overlay_storage, which UEFI VMs therefore require. The firmware writes its variable store and PVE keeps TPM state only in raw volumes, so a read-only map cannot be used. The protected backup is never written. - The PBS client checks every chunk while restoring. The report lists each
image under
source_diskswith its archive and SHA-256, and underrestored_hardware.firmware_diskswith its options, for exampleefidisk0 efitype=4m,pre-enrolled-keys=1andtpmstate0 version=v2.0. The images are always fully restored, test mode included. - Only the options that describe the image are carried over:
efitype,pre-enrolled-keysandms-certforefidisk0, andversionfortpmstate0. - A plan that leaves
biosempty restores a UEFI backup under OVMF. A plan that setsbios: seabiosfor a UEFI backup is refused. - A firmware disk that was excluded from the backup (
backup=0) fails closed, because the source firmware state cannot be reproduced. - The copies hold the source's firmware secrets: Secure Boot variables, and
TPM-sealed keys such as a BitLocker protector. They are
0600, journaled like every overlay, and removed with the sandbox VM, including by crash recovery. - Data disk overlays carry
backup=0, but PVE has no such flag forefidisk0andtpmstate0. A cluster backup job covering all VMs that runs while a validation is in progress therefore copies the sandbox VM's firmware state (its TPM secrets included) into a new backup group, under that job's retention and permissions. Exclude the sandbox VMIDs from such jobs (--exclude, or pool-based selection) or schedule validations outside the backup window.
Copy-before-boot
By default (copy_before_boot: auto), when a tier's disks exceed the cache budget
where reading directly from PBS would cause parts of the guest boot to miss service
start deadlines, CertiStack copies the disks to local overlay storage during the
integrity scan so the sandbox boots from local copies as a full restore does.
auto(default): copies when disks would not all stay cached until boot, provided overlay storage is on a disk separate from the cluster database and has sufficient free space for the disks plus a 5 GiB reserve.always: unconditionally copies disks or fails the VM before boot.never: boots directly from the read-only PBS mapped devices through COW overlays.
storage.max_copy_gib (default: 0, uncapped) limits the cumulative size of copies
made in a single run. Reports record copied_before_boot and each disk's boot_source.
COW-only guest network recovery
Some static guests bind their production connection to a source interface name or PCI path. That binding can change in an isolated recovery VM even when its MAC is preserved. For a VM with Stage 2 wire probes, CertiStack automatically creates one temporary connection profile on the boot disk's disposable COW overlay before boot. You normally need this stanza only for an optional address, legacy interface name, or failure diagnostics:
network_recovery:
mode: auto
# Optional when all Stage 2 probes already target one address.
sandbox_ip: "192.0.2.10"
# Optional compatibility fallback for a legacy guest profile bound to a
# predictable Linux device name; writes a MAC-bound COW-only .link file.
interface_name: "eth0"
# Optional: preserve useful guest-side diagnostics when a wire probe fails.
capture_diagnostics: true
auto detects legacy RHEL/Rocky ifcfg, NetworkManager, Netplan,
systemd-networkd, or Debian-style ifupdown and binds the profile to the
preserved MAC. On EL7, whose NetworkManager ignores keyfile profiles unless
configured to load them, it uses a MAC-bound ifcfg profile. With ifupdown
(networking.service and /etc/network/interfaces), which cannot match a
MAC, it replaces the overlay's /etc/network/interfaces with a static
stanza for the guest's configured interface and keeps the original beside it
as interfaces.certistack-original. Bridges, bonds, VLANs and tunnels are set
aside in favor of the physical interface under them. When the guest names
several physical interfaces and the sandbox VM has only the one NIC (for
example cloud-init's enp6s18 beside an image's stale eth0), each name gets
an allow-hotplug stanza and only the one that exists comes up; with several
NICs (attach_all_nics) it asks for interface_name. It assigns only
the isolated address; it creates no gateway, default route, DNS setting, SNAT,
or physical uplink. The source VM, PBS snapshot, and read-only mapped image
are never changed. The signed VM record lists the chosen strategy, address,
and generated guest path.
When a guest has firewalld installed, the same disposable overlay receives a copy of that guest's effective default zone with source-restricted rules added: the sandbox probe address may reach the configured Stage 2 probe ports and protocols. Everything the guest's own zone already allowed (services, ports, rich rules, target) is kept. This makes an otherwise-correct restored service testable even when its own firewall would refuse the probe, without opening a production firewall or mutating the backup. The change disappears with the COW overlay during cleanup.
Restored guests on the same isolated VNet therefore reach each other exactly as far as their own firewalls allow, as they do in production. A tier that reports on its dependencies (for example a readiness endpoint that checks its peers) can be asserted by a worker probe when the tiers are restored in dependency order. CertiStack adds no rule for guest-to-guest traffic. Routed multi-VLAN topology and application-transaction evidence remain outside the supported boundary until a separate, explicitly qualified isolated-network capability exists.
Set capture_diagnostics: true for a compatibility investigation. On a wire
probe failure, CertiStack records the best safe guest-side evidence available
(interfaces, routes, listening sockets, and relevant service state) in the
signed report. Guest-agent diagnostics are optional: an unavailable or
restricted guest agent is reported as such and does not turn a passing wire
probe into a failure.
interface_name is optional. Use it only for a confirmed legacy profile that
is tied to an interface name (for example ifcfg-eth0); CertiStack adds a
MAC-specific .link file alongside the temporary recovery profile so the
guest sees that name in the sandbox. It never renames a source interface.
Only net0 is attached by default. Set attach_all_nics: true to attach the
source's other NICs (net1 and up) as well: each keeps its MAC and model and
is attached, like net0, only to the isolated VNet with the PVE firewall on.
Their bridges, VLAN tags and firewall settings are never replayed. Use it when
a service binds to a secondary interface's address and would not start
without it. Only net0 gets a recovery address; the other NICs keep the
guest's own configuration, and a guest that waits for DHCP on them boots more
slowly, because the sandbox has no DHCP. Those NICs bring up their
production static addresses on the sandbox VNet, so use the option only when
none of them falls inside network.ip_range: an address there could answer
another VM's probe or clash with probe_ip. The report lists attached NICs
under restored_hardware.additional_nics and the ones left out under
omitted_nics.
To probe a service on another NIC's address, set sandbox_ip to net0's
address and point the probe at the other NIC's address. That address is the
guest's own, so it must be inside network.ip_range to be reachable. It
must also be unique in the plan: validation refuses an address another VM
uses. Whether the guest's firewall admits the probe on that interface
depends on the guest's configuration of it.
network_recovery:
sandbox_ip: "10.0.70.75" # net0, configured by network recovery
attach_all_nics: true
probes:
- {type: http, target_ip: "10.0.70.75", port: 80, path: /healthz, expected_status: 200, timeout_sec: 120, description: "nginx on net0"}
- {type: tcp, target_ip: "10.0.71.202", port: 8080, timeout_sec: 120, description: "service bound to net1's address"}
Omitted network_recovery means auto; CertiStack infers the one guest
address from that VM's Stage 2 probe targets and detects the supported Linux
network backend. Use preserve explicitly to opt out. A plan with wire probes
for more than one guest address fails before recovery rather than silently
guessing which guest identity to rewrite.
Windows and other non-Linux guests
Guest network recovery rewrites Linux network configuration on the disposable
overlay. It cannot adapt a Windows guest, or any guest whose network
configuration it does not recognise, and a VM with wire probes and the default
mode then fails with an unsupported guest network layout error that names
network_recovery.mode: preserve. For such a guest:
- set
network_recovery.mode: preserve, so the guest keeps its own address on the isolated VNet; - make
network.ip_rangethe subnet the guest already uses, and each wire probe'starget_ipthat guest's own address (a disconnected L2 domain can reuse production addressing; see the isolation attestation in the supported recovery contract); - do not set
sandbox_ip, whichpreserveforbids.
Guest-agent probes (qga with sc query <service>) need no guest network at
all. The Windows domain example puts both
to work.
use_backup_config: true remains the recommended path: it discovers all
backed-up disks, safe hardware settings, and the preserved net0 identity.
When full inheritance is intentionally disabled, a run still reads what a
restore cannot do without (firmware, boot order, machine type, SCSI
controller, guest OS type, the backed-up disk list; see "Source disk
layout"), and an automatic wire-probe run also reads the net0 MAC/NIC
identity for the disposable profile. It does not import compute settings,
source network topology, or source tags.
The temporary node worker needs guestfish from libguestfs-tools to make
this overlay-only change. The appliance image/Enterprise connector should
include that dependency. Community reports a clear preflight-style error if it
is unavailable; it never falls back to changing a source guest or attaching it
to production networking.
The signed report separates this offline COW guest-network preparation from mapping and VM boot time, both per VM and in the RTO timing breakdown. This makes guest-layout compatibility costs visible without weakening isolation. "Boot duration" begins at the sandbox VM start request and ends at Stage 1 hypervisor readiness; it does not include storage mapping, guestfs work, or full mapped-image integrity verification.
For a passing validation, rto_seconds and phase_timings.total_rto_sec
measure time from the run start through the last verified recovery probe;
teardown is excluded and reported separately as teardown_sec. A failed
validation has no achieved recovery point, so those RTO fields measure elapsed
time through terminal failure, including teardown when cleanup runs before the
result is finalized. Compare failed-run cleanup duration using teardown_sec,
not by treating its RTO as a successful recovery objective.
3. Deterministic Probing Specifications
Type: tcp (Stage 2: Wire-Level TCP Handshake)
Executes an active TCP 3-way handshake from the node worker's temporary VNet address against the guest IP and port. Stage 2 TCP/HTTP/DNS/LDAP evidence proves worker-to-guest reachability, not an in-guest client path; each signed result records its execution origin. Use a QGA probe when the assertion must execute inside the guest.
- type: "tcp"
target_ip: "192.0.2.10"
port: 5432
timeout_sec: 15
description: "Verify PostgreSQL socket listener on port 5432"
Type: http (Stage 2: Wire-Level HTTP/HTTPS Health Endpoint)
Transmits an HTTP GET request, verifies TLS handshake, and asserts response status code.
- type: "http"
target_ip: "192.0.2.30"
port: 443
path: "/healthz"
tls: true
server_name: "service.example.com" # optional: TLS SNI and HTTP Host header
expected_status: 200
timeout_sec: 20
description: "Verify Web App HTTPS healthz endpoint returns 200 OK"
HTTPS probes always verify the certificate chain and hostname. The recovered
service certificate must chain to a CA trusted by the node worker and its SAN
must include server_name (or the target IP when no server name is supplied).
insecure_skip_verify and tls_skip_verify are rejected deliberately: a
successful recovery assertion must not be based on an unauthenticated TLS
endpoint. For an isolated lab, install the lab CA on the node worker rather
than disabling verification.
Type: dns (Stage 2: Wire-Level DNS Resolution)
Sends one DNS query directly to target_ip over UDP (re-sent until the
timeout) or TCP. The worker's own /etc/hosts and resolver configuration are
never consulted.
- With
domain, the probe passes only when the server answersNOERRORwith at least one A or AAAA record for that name.NXDOMAIN,REFUSED,SERVFAIL, or an empty answer fails, so a restored DNS server that lost its zone is detected. - Without
domain, it is a listener check: the server must return any well-formed reply to a rootNSquery.
- type: "dns"
target_ip: "192.0.2.10"
port: 53
domain: "healthy.example.com"
transport: "tcp" # optional; defaults to udp
timeout_sec: 10
description: "Verify internal DNS listener responds to queries"
With record_type: srv, the probe asks for SRV records instead and passes
only when the answer holds at least one. The report lists each target, port,
priority and weight. For a domain controller, the DC locator record
_ldap._tcp.dc._msdcs.<domain> proves the restored DNS server loaded the
AD-integrated zone, which it reads from the directory, and still carries the
records clients use to find a DC. The records were written before the
backup, so they do not show that Netlogon re-registered after the restore.
- type: "dns"
target_ip: "192.0.2.10"
port: 53
domain: "_ldap._tcp.dc._msdcs.example.com"
record_type: "srv"
timeout_sec: 60
description: "Verify the DC locator records are served"
Type: ldap (Stage 2: Wire-Level Active Directory / Directory Probe)
Performs an LDAP v3 anonymous bind over TCP port 389. When base_dn is set,
it also performs an anonymous base-object search for that DN.
Use an explicit DN when it is a business assertion that must exist. CertiStack
does not silently weaken an explicit DN if the restored directory advertises a
different naming context. For a portable service-recovery check across
heterogeneous customer directories, set base_dn: auto: CertiStack reads the
server's Root DSE and requires a successful base-object search of one of its
advertised naming contexts. Omit base_dn only when an LDAP v3 listener/bind
handshake alone is the intended assertion.
- type: "ldap"
target_ip: "192.0.2.10"
port: 389
base_dn: "dc=example,dc=com" # explicit assertion; requires anonymous search access
timeout_sec: 15
description: "Verify Active Directory domain controller LDAP listener"
- type: "ldap"
target_ip: "192.0.2.10"
port: 389
base_dn: "auto" # portable Root-DSE discovery plus verified base-object search
timeout_sec: 15
description: "Verify any restored directory naming context is readable"
Active Directory. A domain controller refuses anonymous searches by
default, so base_dn (explicit or auto) fails against it, and omitting
base_dn only proves that something answers a bind. Set naming_context
instead: CertiStack binds anonymously, reads the Root DSE (the one entry AD
lets anonymous clients read), and passes only if the directory advertises
that naming context, as defaultNamingContext or in namingContexts. No
directory search is sent and the DC's policy stays as it is. The directory
service answers the Root DSE itself, so a DC whose database did not open
fails even though port 389 accepts connections. The comparison ignores case
and the spaces around , and =. The report records the DC's
dnsHostName and isSynchronized. isSynchronized is informational: a DC
restored without its replication partners may report FALSE while serving
its directory correctly. naming_context works for OpenLDAP and 389-ds too
(their Root DSE lists namingContexts), and cannot be combined with
base_dn.
- type: "ldap"
target_ip: "192.0.2.10"
port: 389
naming_context: "DC=example,DC=com" # Root DSE assertion; no directory search
timeout_sec: 60
description: "Verify the restored domain controller serves example.com"
Type: mssql (Stage 2: Wire-Level SQL Server Protocol Probe)
Sends a TDS PRELOGIN, the first message of every SQL Server connection,
and passes only on a well-formed PRELOGIN response. No login is attempted
and no credentials are needed, so the probe proves that the SQL Server
engine answers its protocol, not merely that the port is open. The report
records the server's version and encryption setting. A named instance must
listen on a static port, because the probe does not query SQL Browser. A
server configured for TDS 8.0 strict encryption expects TLS before any TDS
message, and the probe reports that instead of passing.
- type: "mssql"
target_ip: "192.0.2.20"
port: 1433
timeout_sec: 120
description: "Verify SQL Server answers TDS on the restored guest"
Type: smb (Stage 2: Wire-Level SMB Protocol Probe)
Sends an SMB2 NEGOTIATE over direct TCP and passes only on a successful
NEGOTIATE response for one of the offered dialects (SMB 2.0.2 through
3.0.2). The report records the dialect and whether the server requires
signing. No session is set up and no share is opened, so the probe proves
the SMB server answers, not that a particular share or its data survived.
For that, pair it with an application check that reads the data.
- type: "smb"
target_ip: "192.0.2.21"
port: 445
timeout_sec: 60
description: "Verify the file server answers SMB on the restored guest"
Type: qga (Stage 3: In-Guest Inspection via QEMU Guest Agent)
Uses a constrained QGA health-check vocabulary inside the guest operating
system. Plan-authored commands are limited to hostname, guest-get-host-name,
ping, guest-ping, systemctl is-active <service>, and
sc query <service>; arbitrary shell commands and environment interpolation
are rejected before execution. Guest output lands in the signed report, so
the vocabulary stays read-only and narrow.
- type: "qga"
command: "systemctl is-active postgresql"
expected_output: "active"
require_qga: true # make a missing/unresponsive guest agent a failure
timeout_sec: 30
description: "Verify PostgreSQL service unit state is active inside guest"
On Windows guests, sc query <service> checks a service by its key name
(for example W3SVC, LanmanServer, or MSSQL$SQLEXPRESS). CertiStack runs
sc.exe directly with the literal arguments query and the name, never
through a shell, and passes only when the service's STATE is 4 RUNNING.
A stopped, pending, or missing service fails, and the report quotes the
state or sc.exe's error. expected_output is optional and adds a substring
check. The name may contain letters, digits, and $ _ . -. The Windows guest
agent must allow guest-exec, which it does by default.
- type: "qga"
command: "sc query MSSQL$SQLEXPRESS"
require_qga: true
timeout_sec: 60
description: "Verify SQL Server Express is running inside the restored guest"
Fallback Invariant: If QGA is not installed or unresponsive in the guest, Stage 3 probes are marked as
SKIPPEDand pass/fail criteria falls back strictly to Stage 2 wire verification.
Set require_qga: true for a test whose assertion is the guest-agent channel
itself. In that case an unavailable QGA channel is a probe failure instead of a
skipped optional inspection.
Type: qmp (Stage 1: Hypervisor State)
Connects to the sandbox VM's QMP socket and passes only when query-status
reports the VM running. It proves the VM is powered on, not that its guest
operating system or any service works; use it for a VM that has nothing else
to probe, such as one without a guest agent. The alias qmp_status is
accepted. The result records its origin as hypervisor_qmp.
Type: screendump
Captures the VM's display through QMP and passes when the capture succeeds. It asserts nothing about what the screen shows; it is for the record, in the same way as the forensic capture CertiStack takes when any probe fails.
4. Literal Plans and Protected Runtime Environment
Plans are literal documents. CertiStack does not expand ${VAR_NAME} or
shell syntax in a plan: a plan can be submitted through an API, retained as
signed evidence, and inspected by another operator, so process-environment
interpolation would risk credential disclosure and make its effective content
ambiguous.
Keep connection credentials in a protected allowlisted environment file passed
with --env-file; its values are parsed literally, never shell-sourced. Keep
plan identity and controller-local signing-key paths explicit:
plan_id: "acme-erp-dr"
compliance:
sign_key_path: "/secure/certistack/controller-signing.ed25519"
vms:
- vmid: 9001
source_vmid: 100
source_pbs:
storage_id: "pbs"
snapshot: "latest"
If a site needs a family of plans, render them in a controlled build step,
validate each generated file, and retain its SHA-256 in the campaign's
plans.tsv. Do not run a shell template against a secret-bearing environment.
5. Complete field reference
The parser rejects unknown keys, a plan larger than 10 MiB, and a file that
contains more than one YAML document, so every key it accepts is listed here.
An unknown key's error names its line and section, the key you probably
meant and the keys that section accepts, for example line 4: unknown key
"grace" in tiers[]; did you mean "startup_grace_period_sec"?.
An alias is accepted for older plans and normalised to the canonical key;
write the canonical key in new plans. A rejected key parses but fails
validation, so an unimplemented or unsafe setting can never be silently
ignored. Credentials never belong in a plan: pbs_password, tokens, and
similar keys are unknown to the parser and fail to load.
Top level
| Key | Type | Status | Rule |
|---|---|---|---|
version |
string | required | v1; certistack/v1 is an accepted alias |
plan_id |
string | required | matches ^[a-zA-Z0-9_-]+$ |
name |
string | required | non-empty |
environment |
string | optional | free-text label recorded in the report and shown by validate |
use_backup_config |
bool | optional | inherit safe compute settings and all backed-up disks from each resolved snapshot |
compliance |
map | required | see below |
admission_control |
map | required | see below; admission is an accepted alias |
network |
map | required | see below |
storage |
map | optional, discouraged | see below |
tiers |
list | required | at least one tier |
compliance
| Key | Type | Status | Rule |
|---|---|---|---|
frameworks |
list of string | optional | may be empty; each item one of DORA, DORA-Article-12, SOC2, SOC2-Type2-CC7.4, HIPAA, HIPAA-164.308-A7, ISO27001, ISO27001-A8.13, ISO27001-A.12.3, ISO27001-A12.3, CIS_V8 |
retention_days |
int | optional | 0 (the default) to 36500; recorded in the report, never enforced |
sign_key_path |
string | required on the controller | controller-local Ed25519 private key; never sent to the worker; --key overrides it |
rto_target_seconds |
float | optional | >= 0; when set, a run whose measured RTO exceeds it is reported as failed |
admission_control
| Key | Type | Status | Rule |
|---|---|---|---|
max_host_ram_percent |
float | required | > 0 and <= 100 |
max_host_iowait_percent |
float | required | > 0 and <= 100; max_host_io_wait_percent is an accepted alias |
defer_retry_interval_sec |
int | required | > 0 |
max_defer_retries |
int | required | > 0 |
max_concurrent_boots |
int | optional | 0 (unlimited, the default) to 64 |
boot_settle_sec |
int | optional | 0 (default 180) to 3600; only with max_concurrent_boots |
network
| Key | Type | Status | Rule |
|---|---|---|---|
mode |
string | required | simple or vxlan |
lifecycle |
string | optional | ephemeral (default) or preprovisioned |
zone_id, vnet_id |
string | required | PVE SDN identifier: a letter followed by letters and digits |
tag |
int | required for vxlan |
VNI 1..16777215, unused by every existing PVE VNet; optional for simple |
peers |
string | required for vxlan |
comma-separated IP addresses without duplicates; rejected for simple |
ip_range |
string | required | IPv4 CIDR of the isolated subnet, outside loopback, link-local, multicast, and reserved space; subnet_cidr is an accepted alias |
probe_ip |
string | required | address inside ip_range that is neither the network nor the broadcast address |
gateway |
string | rejected | must be empty; a sandbox never has a gateway |
ipam |
string | rejected | not implemented; must be empty |
mtu |
int | rejected | not implemented; must be unset |
storage
Prefer the protected environment file (PBS_REPOSITORY, PBS_FINGERPRINT)
over this stanza: a repository string can contain a token identity and a plan
is retained as signed evidence.
| Key | Type | Status | Rule |
|---|---|---|---|
pbs_repository |
string | optional, discouraged | when set, a VM may omit source_pbs and defaults to vm/<vmid>/latest |
pbs_fingerprint |
string | optional, discouraged | PBS certificate fingerprint |
pbs_namespace |
string | optional | default for every VM's source_pbs.namespace |
scratch_dir |
string | optional | node-local private work directory (guest network recovery, screendumps; also overlays when overlay_storage is unset); the worker rejects shared temporary directories such as /tmp |
overlay_storage |
string | required on stock Proxmox VE | PVE storage ID of a file-based storage (dir, nfs, cifs, cephfs, btrfs) with images content. Overlays are created as <storage>/images/<vmid>/certistack-delta-*.qcow2 and attached by volume ID. The run fails before any mutation if the storage is inactive, lacks images, or has under 5 GiB free, and it refuses an images/<vmid> directory that holds other volumes. --overlay-storage / CERTISTACK_OVERLAY_STORAGE override it |
max_copy_gib |
int | optional | 0 (no cap, the default) to 1048576. The most copy-before-boot may write in one run, all VMs together. auto does not copy a VM whose copies would pass it, and warns; always fails that VM. Use it where the free space is not all yours, such as a scratch disk shared with other tenants |
tiers[]
| Key | Type | Status | Rule |
|---|---|---|---|
level |
int | required | >= 1, unique across the plan, executed in ascending order; tier is an accepted alias |
name |
string | optional | defaults to Tier <level> |
startup_grace_period_sec |
int | optional | >= 0 |
boot_timeout_sec |
int | optional | >= 0; default 60; PVE start plus QMP readiness |
soak_time_sec |
int | optional | >= 0; hold the tier for this long after its probes pass before starting the next tier; recorded in the report timings |
vms |
list | required | at least one VM |
tiers[].vms[]
| Key | Type | Status | Rule |
|---|---|---|---|
vmid |
int | required | > 0, unique across the plan; the temporary sandbox VM; target_vmid is an accepted alias |
source_vmid |
int | required when snapshot: latest |
the protected source VM |
name |
string | required | non-empty |
mac |
string | optional | six-octet MAC for the isolated NIC; normally inherited with use_backup_config |
nic_model |
string | optional | virtio, e1000, e1000e, igb, ne2k_pci, ne2k_isa, pcnet, rtl8139, vmxnet3 |
machine, ostype, bios, scsihw |
string | optional | explicit PVE values that override inherited ones |
network_recovery |
map | optional | see below |
source_pbs |
map | required unless storage.pbs_repository is set |
see below |
hardware_overrides |
map | optional | cores, sockets, memory_mb, balloon_mb, cpu_weight (ints) |
copy_before_boot |
string | optional | auto (default), always or never. Whether the integrity scan also copies the VM's disks next to its overlays, so the sandbox boots from local copies as a restore does. auto copies when the tier's disks would not all stay cached until the VM boots, unless overlay storage is on the cluster database's disk or lacks room for the disks plus 5 GiB; it warns when it can't. always copies or fails the VM. Reports record copied_before_boot and each disk's boot_source |
require_quiesced_backup |
bool | optional | default false; fail the VM before anything is restored unless its backup is quiesced or powered-off (see Consistency above) |
probes |
list | required | at least one probe |
tiers[].vms[].source_pbs
| Key | Type | Status | Rule |
|---|---|---|---|
storage_id |
string | required | pbs (this run's configured endpoint); other selectors are rejected at run time |
snapshot |
string | required | latest, vm/<vmid>/latest, or vm/<vmid>/<RFC 3339 time> |
drives |
list | required unless use_backup_config |
entries of slot (scsi0..scsi30, virtio0..virtio15, sata0..sata5, ide0..ide3; unique) and archive (PBS archive filename without path separators); every slot must be backed up in the snapshot |
drive |
string | deprecated alias | a single archive, in the slot it names (drive-virtio0) or scsi0 |
namespace |
string | optional | PBS namespace; defaults to storage.pbs_namespace |
tiers[].vms[].network_recovery
| Key | Type | Status | Rule |
|---|---|---|---|
mode |
string | optional | auto (default), preserve, ifcfg, networkmanager, netplan, systemd-networkd, ifupdown |
sandbox_ip |
string | optional | usable IPv4 address in network.ip_range, unique per plan; required when a VM's wire probes target more than one address; forbidden with preserve |
interface_name |
string | optional | Linux interface name (not lo) for a legacy name-bound profile |
capture_diagnostics |
bool | optional | record guest-side diagnostics on a wire-probe failure |
attach_all_nics |
bool | optional | also attach the source's net1+ NICs (MAC and model only) to the isolated VNet |
tiers[].vms[].probes[]
| Key | Type | Status | Rule |
|---|---|---|---|
type |
string | required | tcp, http, dns, ldap, mssql, smb, qga, qmp, screendump; aliases tcp_handshake, http_get, qga_ping, qmp_status |
timeout_sec |
int | required | > 0 (a missing value becomes 10) |
retries |
int | optional | 0..10; 0 keeps automatic polling until the timeout |
description |
string | optional | defaults to <type> verification probe |
target_ip |
string | required for tcp, http, dns, ldap, mssql, smb |
literal IP inside network.ip_range, never probe_ip or the subnet network or broadcast address; a run refuses a target that is an address of the worker host (for example another host address on a preprovisioned VNet) |
url |
string | optional (http) |
http(s)://host[:port]/path; expanded into target_ip, port, path, and tls when target_ip is empty |
port |
int | required for tcp, dns, ldap, mssql, smb |
> 0 |
path, tls, server_name, expected_status |
mixed | http |
expected_status must be > 0; server_name must be a valid DNS name and sets both SNI and the Host header |
tls_skip_verify, insecure_skip_verify |
bool | rejected | any true value fails validation |
domain, transport |
string | dns |
transport is udp (default) or tcp |
record_type |
string | optional (dns) |
a (default: an A or AAAA answer) or srv; srv needs domain |
base_dn |
string | optional (ldap) |
explicit DN, auto, or omitted |
naming_context |
string | optional (ldap) |
a DN the Root DSE must advertise; not with base_dn |
command, expected_output, require_qga |
mixed | qga |
command limited to hostname, guest-get-host-name, ping, guest-ping, systemctl is-active <service>, or sc query <service> (Windows) |