Skip to content

Changelog

All notable changes to CertiStack Community Edition are documented in this file. Enterprise Edition changes are not tracked here.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]

Added

  • Documentation: a command-line reference (every command, flag, exit status and doctor check), a configuration reference (every environment variable, the environment file, and the files CertiStack keeps on the controller and the node), a reports and verification guide with a script that verifies a report without CertiStack, an automation and scheduling guide with a tested wrapper script and systemd units, example plans, and a FAQ. The changelog is now part of the documentation site.
  • examples/three-tier-app.yaml, a Linux database, application and web stack restored in dependency order, and examples/windows-domain.yaml, a Windows domain controller, SQL Server and file server.
  • Getting started explains how to set up SSH access to a node, and the troubleshooting guide covers the node's host lock and the errors that stop a run before it starts.

Fixed

  • Documentation: compliance.frameworks and compliance.retention_days are optional, and the first four admission_control keys are required rather than defaulted. The qmp and screendump probe types, the network_recovery.mode: preserve requirement for Windows guests with wire probes, the OpenSSH client and nft prerequisites, and the fact that verify-report exits 0 for a correctly signed report of a failed validation are now stated.

[0.1.0] - 2026-09-27

First public release of the Community Edition under the GNU Affero General Public License v3.0.

Added

  • storage.max_copy_gib caps what copy-before-boot writes in one run, all VMs together. auto checked only the free space, which on a shared scratch disk is not all the run's to take. A single 48 GiB Windows copy nearly fills the lab's proposed 50 GiB share. Over the cap, auto doesn't copy and warns why, and always fails the VM. plan shows it the same way.
  • admission_control.max_concurrent_boots limits how many sandbox VMs boot at once on the node. A tier starts its VMs one after another, so their first minutes of boot, when each guest writes most to its new overlays, overlapped. On a node whose overlay storage shares the cluster database's disk, that can delay corosync (D36). Before starting a VM, the run waits while that many VMs started within boot_settle_sec (default 180 s), across tiers. plan shows each start it may hold. The run's warning about the shared disk suggests it. It is off unless set.
  • certistack init writes a first plan from what the node and its PBS repository hold, and changes nothing. Without --vm or --all it lists the VMs that have a complete backup. With them it reads each VM's newest backup as a run does (OS type, firmware, guest agent, NIC, disks) and writes a commented plan that validates, with:
  • sandbox VM IDs and SDN IDs that PVE and the other plans in the directory do not use;
  • the node's roomiest usable overlay storage;
  • a first probe per VM that fits its guest: SSH for Linux on its sandbox address, the Event Log service through the agent for Windows, the agent for other guests with one, and the VM running for guests without one. VMs a run could not restore are left out with the reason. From the controller it reads through a temporary worker on the node, as plan does.
  • run, doctor and plan check the PVE token's privileges on every path a run of the plan touches (each VM, the SDN zone and VNet, the overlay storage, the node). They list every missing one at once, against the documented CertiStackRole. A run used to stop at the first missing privilege it met, for example Permission check failed (/sdn/zones/csLab70, SDN.Allocate), and a token missing several needed several attempts. A privilege the run works around is a warning. A test keeps the checked set equal to the role in terraform/main.tf.
  • certistack plan <plan.yaml> shows, step by step, what a run would do on the node, and changes nothing. It lists:
  • what recovery of an interrupted run would delete first;
  • the checks;
  • every resource the run creates, in order: SDN zone, VNet and subnet, probe address and containment table, and per disk the read-only PBS map, the overlay, and any copy made before boot, with the reason;
  • each sandbox VM's exact PVE create parameters;
  • the probes;
  • what teardown deletes, and the evidence it keeps. It also gives the warnings and refusals a run would give, exits non-zero when a run would refuse, and prints JSON with --json. It shares the run's own code for resolving the plan, building the VM configuration and deciding what to copy, so it shows what the run will really do. It sends PVE only reads. From the controller (an SSH host configured, as for run), it stages a temporary worker on the node the same way, has it work the plan out there, and removes it.
  • Copy-before-boot. When a VM's disks would not all stay cached until it boots, the integrity scan also copies them, sparsely, next to their overlays. The overlays are then re-pointed at the copies, and the sandbox boots from local storage as a restore does. It is the same single read of PBS, with the same digest. Large Windows servers no longer boot partly cold on hosts whose NUMA nodes can't hold their disks. The copy releases the map's cache and its own written pages as it goes, so it doesn't crowd out another VM's cached boot.
  • copy_before_boot per VM: auto (default), always or never.
  • auto never copies onto the cluster database's disk, or without room for the disks plus 5 GiB, and warns when it can't copy.
  • Reports record copied_before_boot and each disk's boot_source.
  • A worker workspace that a controller kept for recovery no longer keeps its credentials for ever. The next run on the node, holding the host lock with a clean journal, removes the PVE/PBS environment, the PBS key and the worker binary from any workspace whose controller has been silent for 15 minutes. It keeps the logs and results, and leaves a note saying when and why. Workspaces still in use or being staged are left alone.
  • require_quiesced_backup: true on a plan VM fails that VM, before anything is restored, unless its backup was quiesced or powered-off. An unknown consistency fails too, because it proves nothing. doctor fails such a VM the same way. Without the setting, a crash-consistent backup is still only a warning.
  • The pre-boot scan spreads its page cache over the host's NUMA nodes. The kernel caches a read on the node of the CPU that issued it and evicts from that node once it is full, even while others are free. So a scan left on one node could cache only that node's free memory. In the lab, a 48 GiB Windows VM lost part of its cache during its own scan on a host with two 63 GiB nodes. The scan now moves, every GiB, to the node with the most free memory, within the CPUs the worker may use. The same VM's cache split 42/58 over both nodes, and nothing was evicted.
  • A warning when the disks a tier has scanned exceed what the node can keep cached: its NUMA nodes' free memory together (one node's, the smallest, where the scan can't spread), less the tier's VM memory. Each sandbox boots from the cache its scan filled. When the cache overflows, the first-scanned disks are evicted, those reads go back to PBS at boot, and a Windows guest's services can miss their start deadline.
  • An unknown key in a plan is reported with its line, its section (for example tiers[]), the key it most likely meant, and the keys that section accepts. It used to name a Go type: field grace not found in type config.Tier.
  • When a probe is refused ("connection refused") on a VM restored without some of its NICs, the diagnosis names the missing NICs and suggests network_recovery.attach_all_nics. A service that also listens on a missing NIC's address can fail to start and stop answering on every port.
  • A warning, in every run and in doctor, when sandbox overlays share a physical disk with the cluster file system's database (/var/lib/pve-cluster). Partitions, LVM and RAID are resolved to their disks, and a ZFS dataset to the disks under the devices its imported pool lists (zpool status). When the disks can't be determined, the run treats them as possibly shared. In the lab, a Windows sandbox's first writes to its overlays on a node's single root HDD delayed corosync ("Token has not been received").
  • doctor with a plan reads each VM's backup as a run would, on a node with proxmox-backup-client, and creates nothing. It reports the snapshot, firmware, restored disks, encryption and backup consistency per VM. It fails each VM a run would refuse (no key or the wrong key, UEFI without overlay_storage, a disk the backup lacks) and warns about a crash-consistent backup. --overlay-storage matches run.
  • Backup consistency evidence. Each VM record says whether its backup was quiesced (the guest agent froze the file systems), crash-consistent, taken powered-off, or unknown, with the reason from the backup's own log (source_consistency, source_consistency_detail), and the run warns about a crash-consistent backup. In the lab this found 41 database VMs whose freeze SELinux refused and a Windows domain controller whose VSS freeze expired before the snapshot. The log is read through the PBS API with the configured credentials and certificate trust; the evidence never fails a run.

Fixed

  • Teardown and recovery left an ephemeral sandbox's VNet and zone behind (D46). PVE refuses to delete a VNet that still has a subnet, and says only "Parameter verification failed", so the fallback that deleted the subnets on an error naming them never ran. The subnets are now always deleted first, then the VNet, then the zone. And when nothing could be deleted, teardown no longer applies the unchanged configuration, which only reloaded every node's network and turned forwarding back on for the leftover bridge: it releases the SDN lock and reports the failure.
  • The default ephemeral sandbox network failed on stock PVE 9 (D44): the run refused its own VNet because IPv4 forwarding was on. PVE's ifupdown2 turns forwarding on for every SDN bridge at each network reload. And the run went on as soon as PVE's SDN apply task ended, although that task only forks each node's network reload, one node after another, without waiting for them: 78 s on a 12-node cluster, against a 90 s limit that a larger or slower cluster exceeds (D45). The run no longer waits for that task. It waits for its own node: the VNet's bridge up (and, at teardown, down) and ifupdown2's lock released (ifupdown2 runs as "python3", so the process name cannot tell), then turns forwarding off on its own bridge within a second or two of the node's reload; the containment table drops forwarded packets as well. A failed apply task is reported when the node's bridge does not settle. The zone is created for the run's node only. A pre-provisioned VNet that forwards is still refused, with what to do.
  • An ephemeral sandbox network could apply an administrator's unfinished SDN change. Applying SDN in PVE is all or nothing, on every node, and a run applied the cluster's whole pending configuration along with its own zone and VNet. A run now creates and removes them under PVE's SDN configuration lock (PVE 9): it is refused, before it creates anything, while the cluster has unapplied SDN changes, and the message names them. plan reports it too. At the end of a run over such a change, the sandbox is deleted but the deletion is not applied, and the report says so. The lock is held for seconds; a run that dies holding it leaves its token in the journal, and recovery releases it.
  • Automatic network recovery missed EL7 on a CentOS 7 template (D43): it wrote a NetworkManager keyfile profile, which EL7's NetworkManager 1.18 does not read, so the guest kept its own address and the probes found no one. CentOS 7's /etc/os-release is a symlink, which the EL7 check read as absent, and a plan naming the interface skipped the check. The check now follows symlinks inside the guest, falls back to /usr/lib/os-release, and runs whether or not the interface is named. EL7 gets the MAC-bound ifcfg profile it reads.
  • A plan whose sandbox VMs need more memory than the node can ever hold within max_host_ram_percent is refused at once, by plan (✗, exit 1), doctor and the run, with what it needs, what the node can hold, and what to do: split the plan, or use a node with more memory. Before, plan said "not admitted now; the run would wait and retry", and the run deferred ten times before failing. In the lab, an 84-VM tier needing 249 GiB was refused this way on a 126 GiB node.
  • The controller did not notice a node that crashed or rebooted under a run. The worker command is silent through a grace or soak wait, and without ssh keepalives only TCP keepalive would have noticed, after about two hours. A blade reseated in the lab went unnoticed for 9.5 minutes, until it came back. Now:
  • Every ssh and scp to the node gives up after 60 s without an answer, and after 20 s when connecting.
  • The controller says it lost contact, then retries the node through the 20-minute cleanup window and cleans up as soon as the node answers.
  • It says so when the node rebooted during the run (its boot ID changed).
  • If the node never answers, it says what happens next and prints the exact recover command to run on the node.
  • Lease renewals that keep failing are reported while the run is still going, and so is the return of contact. Before, they were ignored.
  • The final error ends with ssh's reason. Before, it carried the worker's whole stderr, 10 KB in one line.
  • What the controller's own recovery does is shown in the run's output, and the node keeps it in <state-dir>/logs/<run>.log, whoever started the recovery. Before, only a recovery the worker's supervisor ran left a log on the node.
  • Release and CI builds use Go 1.26.8, including the digest-pinned controller builder, and check reachable vulnerabilities with the build toolchain.
  • The signed release checksum manifest authenticates CONTAINER-IMAGE.txt as well as the CLI binary. Container and appliance instructions verify it before using the image digest.
  • Controller PBS encryption keys must be regular, owner-only source files. Staging reads the checked file descriptor and rejects symlinks, including replacement during open on Linux.
  • The explicit disposable-lab TLS acknowledgement now passes through controller environment files and the protected worker environment. An omitted acknowledgement clears any inherited value.
  • Storage documentation describes automatic full-disk copies, temporary firmware copies, overlay growth, and the default uncapped copy policy.
  • plan said a run would refuse when an unfinished run's sandbox VM still held the plan's VMID ("already exists in the PVE cluster"). The run recovers that VM first and goes on. plan now counts what its "before the run" recovery deletes, the unfinished run's VMs, zone and VNet, as gone, as the run finds them (S15-07b). It also lists the probe address's nftables containment table, which recovery removes with the address.
  • A run stopped by a signal now says which, and whether the node was shutting down. After a node reboot, the report and the controller used to say only "validation failed: context canceled", so an operator couldn't tell a node shutdown from a Ctrl-C or a bug (S15-07). The reason now reads, for example, "interrupted: the node began shutting down (SIGTERM from systemd); the run was at: …". The controller reports an interrupted run, exit status 130, not a failed validation.
  • A sandbox VM that QEMU stopped went unnoticed until its probes timed out as network errors. The case found was a guest stopped on a write to a full overlay storage (S15-06e): PVE's status still says "running" there; only qmpstatus says io-error. The run spent its whole grace period and probe timeouts, then reported "no route to host", and the report never named the full storage. The run now checks every 10 s, while it waits, probes and soaks, that its VMs have not stopped for good. A VM that has fails at once, with the reason:
  • a disk write error, with the room left where its overlays are;
  • a QEMU internal error;
  • a guest OS crash;
  • a shutdown.
  • plan refused every copy_before_boot: always VM ("the free space where the copies go is unknown"), where the run copies. It measured the sandbox VM's image directory, which only the run creates, so its free space and disks came back unknown. auto could also plan a copy onto the cluster database's disk that the run would refuse. plan now measures the directory's storage (its nearest existing parent, never above the storage itself), less the copies of the VMs before it, as the run sees it.
  • Recovery could hang the node's management plane. A map daemon killed while one of its threads read its own file (the loop's partition scan) never exits, so its FUSE connection stays open. Every open of the loop then waits in the kernel for ever: recovery's losetup -d, udev, and PVE's vgs, which froze the node's storage status for minutes in the lab (S15-02f). Before it detaches a map it owns whose daemon is gone, recovery now aborts the map's FUSE connection, found through the mount table. It runs the unmap and the detach under a 20 s watchdog, aborts the connection when either hangs, and reports a detach that is still stuck instead of waiting for it. doctor and preflight never open a loop with reads in flight. They name FUSE connections with requests waiting and no mount, with the abort command to run before losetup -d.
  • An ephemeral sandbox, the default, could not start on PVE 9. Before creating its zone and VNet, a run checked that neither existed by reading each one. PVE answers a read of a missing zone with HTTP 500 ("sdn zone '' does not exist"), not 404, so every free ID looked like an error and the run and plan stopped at the ownership check. The lab found it through init: its sandboxes are pre-provisioned, which skips the check. The check now reads PVE's zone and VNet lists. Teardown of a zone or VNet that is already gone accepts the same HTTP 500, but only when it names that object.
  • A run with a pre-provisioned sandbox read its zone at /cluster/sdn/zones/<zone>, which PVE allows only with SDN.Allocate. A token without it passed the privilege check and then failed at that read (S14-18b). The run now reads the zone and VNet from PVE's lists, which need only SDN.Audit, and the privilege check asks for SDN.Audit on the zone.
  • plan showed the overlay storage check as passed where the run warns: overlay storage on the cluster database's disk (D36), and a PVE storage status that did not answer. Both are now a warning on that check and in the plan's warnings.
  • Recovery after a worker was killed while it mapped a disk could report its journal clean and still leave the map behind (S15-02d, S15-02e). A map daemon stopped before it wrote its pid file left a loop device or a dead FUSE mount, and recovery found a map only through its pid file. Recovery now also finds the map by its name in /proc/self/mountinfo and /run/pbs-loopdev. It waits for killed map processes to exit, and detaches a dead mount with umount2(MNT_DETACH), which works where umount -l gets ENOTCONN. A mount it cannot remove is reported, and recovery says incomplete, not clean. doctor and the run's preflight also warn (stale_loops) about a read-only loop attached to a FUSE file that is no longer mounted, which nothing ties to a run, and give the losetup -d command that releases it.
  • After a node worker was killed, what its supervisor's recovery did went unseen. The supervisor stopped streaming to the controller first, and the log was deleted with the workspace. Recovery now reports every action as a progress event ([Recovery] Destroyed sandbox VM …, Stopping the PBS map process …, Releasing /dev/loopN …). The supervisor streams them to the controller until recovery ends, and keeps the log in <state-dir>/logs/<workspace>.log on the node, also when the controller had gone.
  • A run failed at its first check when PVE's status read of a local overlay storage did not answer: "read storage ... status: PVE API error (HTTP 596)". PVE computes every storage's status on the node together, so a slow PBS storage, as during its backup window, stalled it; in the lab for 2 min 19 s. The read now gets 45 s. When PVE doesn't answer, the run checks the storage on the node itself: the directory pvesm path resolves for it must have room. It warns and names the cause. A definite answer from PVE, such as the wrong storage type, still refuses.
  • A worker killed while proxmox-backup-client map ran, or a map that timed out after creating its loop device, left a read-only loop device and map daemon that no journal named. Each map is now journaled before it is made. Recovery, and the run itself when a map fails, releases the loop device only when all of these hold:
  • its PBS map file name ends with that snapshot and archive;
  • its daemon started within the map's window;
  • no journal already tracks it. It first stops the map's own client process, since the daemon it forks outlives a killed worker and can create its loop device after recovery has looked. It sends SIGINT, as proxmox-backup-client unmap does, and SIGKILL only after 20 s. The process must match by its arguments and start time. It then removes the stale FUSE mount and pid file a killed daemon leaves behind, which the client's own cleanup would trip over.
  • Sandbox overlay disks use cache=unsafe and are created with lazy refcounts. Under PVE's default (cache=none), QEMU also read the backup through its qcow2 backing with O_DIRECT, so every read bypassed the pages the pre-boot integrity scan had cached and went back through the PBS map. Restored Windows Server 2016 and later could then miss the service manager's 30 s start deadline: the guest agent or a database service stayed stopped. The sandbox now boots from the cache the scan fills. In the lab, a Windows Server 2016 guest that had failed this way started its agent at the first check, with the node's I/O wait under 1% during boot. The overlays are disposable, so skipping the guest's flushes loses nothing. Runs with the lab-only --skip-integrity still boot cold.
  • The pre-boot integrity scan releases its own cached copy of each disk as it reads. A mapped disk is a loop device over the PBS map's FUSE file, so a buffered read cached every block twice. Only the FUSE file's copy survives the scan, and it is the copy the sandbox boots from. The other copy doubled the scan's memory use and could push the boot's copy out on a node with less free memory than twice the VM's disks.
  • After the node worker was killed (SIGKILL, or with its SSH session), the controller refused to clean up: "recorded node worker PID no longer identifies this workspace binary" (exit 2). The sandbox VM, PBS maps and probe address then stayed live until the node's next run. A worker that no longer runs, or whose PID now belongs to another program, is now left alone and its journal is recovered at once. Recovery takes the host lock, which a live worker would hold, and acts only on the identities the journal recorded. If the controller died too, nothing recovered the journal before the node's next run. The worker's supervisor on the node now recovers it itself when the worker dies by a signal without a result.
  • A Windows sandbox VM could run a newer machine version than a restore of its backup. Proxmox VE starts a Windows guest with an unversioned machine type on the version of the QEMU that created it (5.1 when that QEMU is older than 9.1), which it reads from the VM's meta line. PVE writes the sandbox VM's own meta and refuses to change it. So when a restore would start on another version the node can run, the sandbox VM is now pinned to it, and the run says so. Where both agree, it keeps the backup's unversioned setting.
  • The full mapped-image integrity scan ran while the sandbox VM booted, reading every byte of every disk through the same PBS map as the guest's first reads. The guest's boot starved, and Windows services can time out at start. The scan now completes before the VM starts, as a restore has all its data before it boots, and a damaged backup fails before anything boots from it. The run takes no longer: the guest's boot already fell inside the startup grace period that followed the scan.
  • The failed-VM diagnostic listed only failed units. A data mount skipped after its device timed out (nofail), and the services left inactive by it, did not show at all. The diagnostic now lists mounts that are not active and the boot's dependency, device-timeout and mount failures. The live (guest agent) and offline (overlay) diagnostics both do.
  • The documented CertiStackRole lacked VM.Config.CPU and VM.Config.Memory, which every VM create needs. A token with exactly that role could not run at all. The role now lists them, and VM.Config.CDROM for the CD-ROM drives below.
  • The sandbox VM had no CD-ROM drives where its source has them (a cloud-init drive, an ISO), which a restore keeps, so the guest saw a storage controller disappear. They are now attached empty, never in place of a restored disk. A token without VM.Config.CDROM falls back to root's local qm on the node, or creates the VM without them and warns. The signed record lists the drives the VM actually has (restored_hardware.empty_cdroms).
  • A passphrase-protected PBS key without PBS_ENCRYPTION_PASSWORD, or with the wrong one, failed inside proxmox-backup-client, whose output CertiStack withholds, so the run showed only an exit status. A key whose file says it needs a passphrase is now refused before anything is staged or restored when none (or one shorter than PBS accepts) is set. A wrong passphrase is named in CertiStack's own words; the client's output stays withheld.
  • A clean run could end "cleanup unverified" (exit 2) and keep its worker workspace, credentials included, when another run started on the same node between the worker finishing and the controller's own check. That check needed the host lock, and the node's journal might already belong to the new run. The worker now records the run under an ID the controller chose. The controller reads the journal without the lock and accepts it when it is this run's and clean, absent, or a later run's: a run only replaces a clean journal, after recovering it. Only this run's own unclean journal is recovered, and a busy host lock is waited out instead of failing. With --ssh-sudo, the root-only state directory is now checked through sudo. Before, every file in it looked absent.
  • The sandbox VM got a new SMBIOS identity, which a restore of the same backup does not: qmrestore keeps smbios1 (it regenerates the UUID only with --unique). The backup's SMBIOS UUID and vendor strings are now carried, each validated as Proxmox VE accepts them, and the signed record lists the UUID (restored_hardware.smbios_uuid). The VM generation ID is still new, as after any restore.
  • A Windows sandbox VM ran a newer machine version than a restore of the same backup would. Proxmox VE pins a new Windows VM's machine version on create (the backup's q35 became pc-q35-11.0+pve2), while qmrestore keeps the backup's unversioned setting, which the same nodes resolve to pc-q35-11.0+pve0. The backup's setting is now restored after create, and the signed record carries the version the VM ran (restored_hardware.running_machine).
  • The sandbox VM always had a memory balloon device, even when its source had none (balloon: 0, common on Windows guests), so the restored guest saw hardware it never had. The device now follows the backup unless the plan sets hardware_overrides.balloon_mb, and the signed record says so (restored_hardware.no_balloon).
  • One transient PVE API failure failed the whole run: pveproxy answers HTTP 596 when a cluster node busy with clones or backups does not respond in time, and the same read succeeds a second later. Reads (GET) are now retried up to four times over about 17 seconds on HTTP 502-504 and 595-597 and on dropped connections, and every retry is logged. Other requests are retried only when they provably never arrived (a refused connection, or 595); timeouts are not retried. Errors for 595-597 now say what failed instead of an empty message.
  • Network recovery refused a Debian ifupdown guest whose configuration names more than one interface, which every cloud-init Debian image does (cloud-init's enp6s18 beside the image's stale eth0), and asked for network_recovery.interface_name. When the sandbox VM has only the one NIC, each name now gets an allow-hotplug recovery stanza and only the one that exists comes up. Bridges, bonds, VLANs and tunnels are set aside in favor of the physical interface under them. VMs restored with several NICs still need interface_name.
  • sc query checks copied sc.exe's state text into the report unbounded; a hostile guest could put about 600 KB, or terminal escape sequences, into a signed report. The state is now reported as its code and a fixed name. It is found by value, so translated labels ("STATUS" on German Windows) no longer fail a running service. All guest text in QGA results is now stripped of control characters as well as capped.
  • A DNS srv check passed on an SRV record with target ".", which RFC 2782 defines as "service not available". Such records no longer count. Probe details list at most eight answers, so a reply packed with compressed records cannot inflate the report.
  • ldap naming_context failed when a domain controller refused the clear-text anonymous bind (strongAuthRequired or inappropriateAuthentication), although it still serves the Root DSE. The check now reads the Root DSE after such a refusal.
  • proxmox-backup-client ran without the parent-death signal every other helper uses. A restore or map started just before the worker was killed could finish after crash recovery, leaving a TPM-state copy or a loop device behind. PBS commands now die with the worker.
  • A drive archive such as drive-scsi0.img.fidx.fidx was rewritten to a mangled name and passed to PBS; it is refused now.
  • compliance.retention_days is bounded at 36500 (100 years). Larger values are signed into reports and overflowed evidence stores' durations.
  • With CERTISTACK_PBS_KEYFILE set, every run failed: the key was passed to snapshot list and snapshot files, which do not accept it. It also went to unsigned snapshots, whose restores then fail the manifest signature check. The key now goes only to restore and map, and only for encrypted or signed snapshots, so one plan can mix both.
  • SCSI disks got iothread=1 behind every controller. PVE honors it only with virtio-scsi-single, and elsewhere its warning failed the VM start. A PVE task that ends with WARNINGS: N now counts as successful, as it does in PVE.
  • Without a boot order, CertiStack picked the first disk by ide, sata, scsi, virtio. PVE's order is ide, scsi, virtio, sata, so a VM with a SATA data disk booted and network-recovered the wrong disk.
  • A source machine with options (q35,viommu=virtio) refused every plan. It is parsed as PVE's property string now, keeping allowlisted options.
  • A source without a scsihw or ostype line runs on PVE's defaults, lsi and other, and now restores on them rather than on virtio-scsi-single and l26.
  • use_backup_config refused a VM over any unusable net1+ line. Those NICs are now reported as omitted (or refused only if attach_all_nics asks for them), and NIC errors name the real slot.
  • A plan could restore one archive into two slots, or an archive into another slot, and the omitted-disk evidence would be wrong. Each archive must now go to its own slot, once.
  • source_crypt_mode was the strongest mode of any file in the snapshot listing. It is now the weakest mode among the VM's disks, and the new source_signature_verified says whether a key verified the manifest. Disks excluded from backups (backup=0) are listed as excluded_source_disks instead of silently missing.
  • When the controller keeps a worker workspace for recovery, it now deletes the staged PBS encryption key from it; recovery never needs the key.
  • doctor and the node worker's preflight now check what the sandbox needs before anything starts: nft when the plan has wire probes, guestfish when it uses network recovery (with the package to install, or mode: preserve to opt out), and a configured PBS encryption key.
  • A wrong PBS key (another storage's key) failed inside the PBS client with a withheld error. CertiStack now compares the key file's fingerprint with the snapshot's key fingerprint from snapshot list and refuses before any restore, naming both. A file that is not a PBS key file (a paperkey printout, a PEM master key) is refused at startup.
  • A run that missed compliance.rto_target_seconds signed a failed report but exited 0; it now fails like any other validation failure, and the controller never reports success for a worker report that does not record a pass.
  • A QGA probe with expected_output on a guest that disables guest-exec silently fell back to matching the host name, so a check such as systemctl is-active postgresql could pass without checking the service. Only hostname probes use that fallback now.
  • The dns probe passed on NXDOMAIN, and without domain it looked up localhost, which the worker answered from its own /etc/hosts without contacting the guest. It now sends the query directly to the target and, with domain, requires an A or AAAA answer.
  • Multi-tier plans reserved memory for the largest tier only, although every earlier tier keeps running until teardown; admission now reserves the sum.
  • The CLI controller ignored storage.scratch_dir; it now applies after --scratch / CERTISTACK_SCRATCH_DIR and before the built-in default.
  • validate now rejects plans the runtime would refuse: a storage_id other than pbs, mixed PBS namespaces, vmid equal to source_vmid, SDN zone or VNet IDs longer than 8 characters, and http probes without a port. The shipped minimal example used 9-character SDN IDs and was corrected.
  • On stock Proxmox VE the sandbox VM could not be created: the API refuses file paths for an API token, and even root's qm create refuses a disk outside a storage ("unable to associate path ... to any storage"). The new storage.overlay_storage setting (--overlay-storage) places overlays as VM-owned volumes of a file-based storage and attaches them by volume ID with the existing least-privilege token. The storage is validated before any mutation. Without it, the old file-path mode now explains how to fix the failure.
  • Restored Ubuntu 26.04 (netplan) guests started their services about two minutes late, so Stage 2 probes timed out. netplan resolved the MAC-matched recovery profile to the NIC's name at generator time (ens18, after a dracut initrd rename), the guest's own set-name then renamed it back, and systemd-networkd-wait-online waited for a name that no longer existed. For netplan guests the overlay now carries a wait-online drop-in that waits for any online link. The profile also no longer marks the link optional and disables IPv6 autoconfiguration, which the air-gapped sandbox never answers.
  • Network recovery silently did not take effect on two guest families, which booted with their own address (and, on Debian, their own default gateway): Debian-style ifupdown guests (auto picked systemd-networkd because /etc/systemd/network exists) and EL7 guests (NetworkManager 1.18 ignores keyfile profiles unless the keyfile plugin is configured). Auto now detects ifupdown and EL7 and writes a profile those guests actually apply.
  • latest could resolve to an aborted backup (partial files, no index.json.blob manifest), and every restore of that VM then failed with "restore PBS VM configuration: exit status 255". It now resolves to the newest complete snapshot and logs the incomplete ones it skipped; a group with only incomplete snapshots fails with an explicit message.
  • A probe that polled until its timeout reported only the final attempt, which the deadline had cut off ("timed out waiting for guest process"), and hid what the guest had answered on every completed attempt. It now reports the last completed attempt. A failed QGA command also shows its stdout when stderr is empty (systemctl is-active answers inactive on stdout).
  • guestfish failures lost guestfish's own error and always suggested installing libguestfs-tools, so an empty or non-Linux disk ("no operating system was found on this disk") read as a missing dependency. The first guestfish error line is now included, and the install hint appears only when guestfish is missing.
  • Every restore failed with proxmox-backup-client 4.2, which reports the mapped device as "mapped on /dev/loopN" instead of "mapped to"; both wordings are accepted.
  • scripts/run-lab-campaign.sh could never credit a --skip-integrity campaign and did not recognise the engine's runtime admission denial.
  • An http probe whose target answered with the wrong status on every attempt reported only "context deadline exceeded" when the probe deadline ended mid-attempt, hiding the status the target had returned. It now reports the last answer (for example "expected HTTP status 200, got 503 Service Unavailable") and says that later attempts ran out of time. The tcp probe had the same flaw: a stopped listener was reported as "context deadline exceeded" instead of "connection refused".
  • A scheduled backup job that selects all VMs backed up the sandbox VMs of a running validation, copying restored guest data into the backup store under the sandbox VMIDs and locking each VM for minutes. Overlays are now attached with backup=0, so such a job stores only the sandbox VM's configuration. docs/deployment-model.md explains how to exclude sandbox VMs from those jobs.
  • The documented Docker Compose deployment could not start a run: the root filesystem is read-only and nothing pointed the controller's temporary directory at the /tmp/certistack tmpfs. The image, Compose file, and appliance wrapper now set TMPDIR=/tmp/certistack, and CI and the release start a run from the image under the published restrictions (scripts/controller-container-smoke.sh) and require it to reach its first node contact.
  • A worker whose recovery journal could not be written after it assigned the host probe address (and its nft containment table), or after it created the SDN zone or VNet, returned without a teardown step for that resource. Removal is now registered as soon as each resource exists, and the probe address is journaled as a pending allocation before it is assigned.
  • An unreadable signing key or an unwritable report directory was found only after the remote run, when the evidence was about to be removed with the node workspace. The controller checks both before contacting the node, and a report that still cannot be saved is returned and printed between BEGIN/END CERTISTACK REPORT markers instead of being lost.
  • An arm64 controller uploaded itself as the worker to an x86_64 PVE node. The controller now checks the node's architecture before uploading, and releases publish amd64 only; arm64 remains a source build for the offline commands.
  • keygen could leave a new private key without its public half when the public path was refused, and --private k --public k --force replaced the private key with the public one. Both destinations are checked before either is written, the private key is restored if the public write fails, and paths naming one file are refused.
  • scripts/generate-current-coverage-cases_test.sh, run by CI and the release, failed: the generator wrote its --storage-id value into storage_id, which the engine now restricts to pbs. The flag defaults to pbs and rejects anything else.

Changed

  • The supported pair is Proxmox VE 9.x with PBS 4.x, the pair every live campaign ran on. The docs used to list PVE 8.x and PBS 3.x as well; they are not supported until they have their own signed matrix evidence.
  • On firewalld guests, network recovery keeps the guest's own default zone and adds the probe rules to it, instead of replacing the zone with the probe rules only. Restored guests on the same isolated VNet now reach each other as far as their own firewalls allow, as in production. Before, a multi-tier application failed in the sandbox because its Rocky tiers refused their peers even on ports their own zone opened. CertiStack still adds no rule between guests, and the sandbox stays air-gapped.
  • Every plan now reads the backup's firmware settings (bios, efidisk0, tpmstate0), not only plans with use_backup_config. A plan with explicit drives used to restore a UEFI VM under SeaBIOS without its firmware state, which cannot boot. A plan that leaves bios empty now gets OVMF for a UEFI backup, and bios: seabios for a UEFI backup is refused.
  • The source configuration parser stops at the first [section]: snapshot and pending sections describe other states of the VM.
  • A plan that restores only some of a VM's backed-up disks now says so: the disks it leaves out are listed in the VM's omitted_source_disks in the signed report, and the run warns. A plan that lists a disk the backup does not have is refused before anything is created. Before, such a plan restored a partial VM without saying so.
  • The single-disk source_pbs.drive shorthand uses the slot its archive names (drive-virtio0 is virtio0) instead of always scsi0.
  • Plans without use_backup_config now also take the backup's machine type, SCSI controller and guest OS type when they do not set them. They used to get PVE's defaults (i440fx, virtio-scsi-single, l26). A guest installed on q35 or an LSI controller could fail to boot there, which says nothing about the backup.

Added

  • Report schema 1.5: each VM record carries source_vmid and restored_hardware (the virtual hardware the sandbox VM was created with), and phase timings include integrity_verification_sec.
  • network_recovery.mode: ifupdown for Debian-style guests. The overlay's /etc/network/interfaces is replaced by one static stanza for the guest's single configured interface (no gateway, no DNS); the original is kept as interfaces.certistack-original. Guests with several interfaces need network_recovery.interface_name.
  • When a failed VM's guest agent shows the recovered NIC without its recovery address, the run says so ("network recovery did not take effect: guest NIC ... has ..., not the recovery address ...") in the progress output and the signed diagnostic.
  • ldap probes take naming_context: an anonymous Root DSE read that must advertise the given naming context, with no directory search. It asserts an Active Directory domain controller, which refuses anonymous searches by default, instead of settling for "port 389 answers a bind". The report records the DC's dnsHostName and isSynchronized.
  • dns probes take record_type: srv, which requires SRV records for domain (for example the DC locator _ldap._tcp.dc._msdcs.<domain>) and reports each target, port, priority and weight.
  • qga probes accept sc query <service> on Windows guests. sc.exe runs directly with literal arguments (no shell), and the probe passes only when the service's STATE is RUNNING. sc.exe exits 0 for a stopped service too, so the exit code alone is not used. Until now a Windows guest could only be checked with QGA ping and hostname, because in-guest commands ran through /bin/sh -c.
  • mssql and smb probes. mssql sends a TDS PRELOGIN and passes on a well-formed PRELOGIN response, recording the server version and encryption setting. smb sends an SMB2 NEGOTIATE and passes on a successful response, recording the dialect and whether signing is required. Neither logs in, so both prove that the service answers its protocol, where a tcp probe only proves that a port accepts connections.
  • UEFI VMs restore with their firmware state. The UEFI variable store (efidisk0) and TPM state (tpmstate0) are restored from PBS as disposable raw copies on storage.overlay_storage and attached in the source slots. The report records each copy's archive and SHA-256 under source_disks, and its options under the new restored_hardware.firmware_disks. Before, a backup with either disk was refused, so no UEFI VM, which includes most Windows Server 2022/2025 VMs, could be validated.
  • Disks on every bus restore: VirtIO block (virtio0-virtio15), SATA (sata0-sata5) and IDE (ide0-ide3) as well as SCSI, each through the same read-only map and COW overlay. The sandbox VM boots the first restored disk of the source's boot order (recorded as restored_hardware.boot_disk), so a VM that boots from virtio0 or sata0 restores as it runs. Before, such backups were refused, or restored without their non-SCSI disks.
  • network_recovery.attach_all_nics attaches a VM's other NICs (net1 and up) with their MACs and models, only to the isolated VNet, so services bound to a secondary interface's address start. With sandbox_ip set to net0's address, wire probes may also target those NICs' own addresses; every guest address in a plan must be unique. NICs left out are listed in the report's omitted_nics. All PVE NIC models are accepted, including the e1000 variants of VMs migrated from VMware (e1000-82545em), which use_backup_config used to refuse.
  • Encrypted PBS backups restore. Set CERTISTACK_PBS_KEYFILE to the backup encryption key (and PBS_ENCRYPTION_PASSWORD for a password-protected key). The controller stages the key in the run's private worker workspace, and the worker refuses a key that other users can read. Every PBS call passes it, and each VM record states the backup's source_crypt_mode. Before, the key was never passed, so an encrypted backup failed at its first restore with an error that did not say why. Now a missing key stops the run before anything is restored, with a message naming the setting.

Security

  • A signed report with a duplicated JSON key could carry values its signature never covered: canonicalization keeps the last duplicate, while decoding merged the earlier one, so a failed probe could be turned into a skipped one. Verification now rejects duplicate keys at every level and decodes the result from the exact bytes whose signature it checked.
  • Wire probes (tcp, http, dns, ldap) used ordinary host sockets, so a plan with a loopback sandbox range passed against a listener on the worker host with no guest present. Probes are now sent from the host probe address and bound to the sandbox VNet interface, and validate rejects a sandbox range that overlaps loopback, link-local, multicast, or reserved space, and probe targets on the subnet's network or broadcast address.
  • With PVE_TLS_INSECURE=true but no lab acknowledgement, node-run recovery and preflight and doctor sent the PVE API token over an unverified connection before the runner refused the run. The PVE client itself now refuses to send any request in that case.
  • A bound wire probe could still pass against the worker itself: a host listener on another address of the sandbox interface (for example on a preprovisioned VNet) answered it with no guest present. A run now refuses any probe target that is an address of the worker host, before it creates anything and again on every connection attempt.
  • When assigning the host probe address failed and removing its nftables containment table also failed, the table stayed installed while the report said all_cleaned and the journal said cleaned. A failed rollback is now reported, the leftover table (or address) is journaled and removed at teardown and by recovery, and the evidence stays unclean until that succeeds. An address already on the interface is refused rather than treated as the run's.
  • A failing http probe signed up to 200 characters of the response body, which can carry a stack trace or application data, into the report. Only the status code is recorded now.

Initial release capabilities

  • Controller CLI: validate, run, verify-report, keygen, doctor, inspect, recover, and version.
  • Controller-to-node execution model: strict-SSH bootstrap of a short-lived node worker under /var/lib/certistack/workers/<uuid>, allowlisted literal environment files, and no persistent service on the hypervisor.
  • Mapped recovery: read-only proxmox-backup-client map of PBS snapshots with per-disk qcow2 COW overlays on private node scratch storage, full mapped-image integrity verification, and multi-SCSI-disk layouts.
  • Backup-configuration inheritance (use_backup_config) for safe compute, firmware, machine, SCSI controller, MAC, and NIC model settings.
  • Isolated PVE SDN sandboxes (simple and vxlan, ephemeral or preprovisioned) with no gateway, no SNAT, run-scoped nft host containment, and refusal to adopt pre-existing zones or VNets.
  • COW-only guest network recovery for ifcfg, NetworkManager, Netplan, and systemd-networkd guests, including a scoped firewalld probe rule; the source VM and backup are never modified.
  • Deterministic probes: QMP readiness, worker-origin TCP, HTTP/HTTPS, DNS, and LDAP wire probes, a constrained QEMU Guest Agent vocabulary, and forensic QMP screendumps on failure.
  • Host admission control with capacity reservation for the largest concurrent tier.
  • Durable run journal, LIFO teardown stack with signal and panic unwinding, and journal-scoped crash recovery with process-identity verification.
  • Ed25519-signed RFC 8785 canonical JSON reports (schema 1.4) with pinned signer verification and machine-readable verification output. The JSON is the evidence; HTML/PDF binder rendering and outbound notifications are Enterprise Edition features and are not part of this repository.
  • Lab campaign runner with resume, fault, and guard modes, plus the coverage-case generator and compatibility-matrix audit tools.
  • Unprivileged controller OCI image, Docker Compose reference, Packer appliance template, and a Terraform example for least-privilege PVE RBAC.
  • Documentation site (MkDocs Material), example plans, contributor license agreement, code of conduct, and security policy.
  • Library packages for products built on the engine: pkg/runner, pkg/recovery, pkg/attest, and pkg/config.

Security

  • Chaos fault-injection hooks compile only with -tags=chaos and are a constant false in every shipped binary.
  • TLS verification bypasses are rejected in plans; PVE_TLS_INSECURE requires an explicit lab acknowledgement and is refused for release evidence.
  • PBS repository strings, credentials, and raw PBS-client output are withheld from logs, reports, and campaign bundles.
  • Release binaries ship with a GPG-signed checksum manifest and SLSA build provenance.