Verify backups end-to-end
Make sure that WAL archiving and base backups reach S3, and that you can restore a Cluster from them. The page includes the queries, the bucket check, and a restore test.
The status fields of the CloudNativePG resources show what the operator expects. The bucket shows what the backup really wrote. A restore shows that you can use the backup. Do the checks on this page after you deploy a new Cluster or migrate an existing Cluster. See Deploy a cluster with the barman-cloud plugin and Migrate from the legacy in-tree backup.
Before you start
Set these shell variables. The commands on this page use them.
CLUSTER="<cluster>" # name of the Cluster
BUCKET="<bucket>" # bucket from destinationPath in the ObjectStore
PRIMARY=$(kubectl get cluster.postgresql.cnpg.io $CLUSTER -o jsonpath='{.status.currentPrimary}')
REPLICA=$(kubectl get pods -l cnpg.io/cluster=$CLUSTER,cnpg.io/instanceRole=replica \
-o jsonpath='{.items[0].metadata.name}')After a failover, another instance is the primary. In that case, set PRIMARY and REPLICA again.
The commands use the current namespace of your kubectl context. Some commands run psql as the user postgres in the database postgres. The postgres container allows this user through peer authentication on the local socket.
The example outputs come from a Cluster with the name pg.
1. Cluster and pods
kubectl get cluster.postgresql.cnpg.io $CLUSTER
kubectl get pods -l cnpg.io/cluster=$CLUSTER -o wideHealthy output:
NAME AGE INSTANCES READY STATUS PRIMARY
pg 10m 3 3 Cluster in healthy state pg-1
NAME READY STATUS RESTARTS AGE ... NODE
pg-1 2/2 Running 0 10m ... kub-d2
pg-2 2/2 Running 0 10m ... kub-d7
pg-3 2/2 Running 0 10m ... kub-d8READY 2/2 means that each pod runs two containers: postgres and the plugin sidecar plugin-barman-cloud. A Cluster with the legacy in-tree backup shows 1/1.
To make sure that the second container is the plugin sidecar, run:
kubectl get pod $PRIMARY -o jsonpath='{.spec.initContainers[*].name}{"\n"}'
# bootstrap-controller plugin-barman-cloud
kubectl get pod $PRIMARY -o jsonpath='{.spec.containers[*].name}{"\n"}'
# postgresThe sidecar is a Kubernetes native sidecar: an init container with restartPolicy: Always. It continues to run after the pod starts.
2. WAL archiving
WAL archiving runs on the primary. Read the archiver counters on the primary:
kubectl exec $PRIMARY -c postgres -- psql -U postgres -d postgres -c \
"SELECT archived_count, failed_count, last_archived_wal, last_archived_time FROM pg_stat_archiver;"Healthy output:
archived_count | failed_count | last_archived_wal | last_archived_time
----------------+--------------+--------------------------+-------------------------------
10 | 0 | 000000010000000000000008 | 2026-10-04 22:15:41.587741+00Write a marker row and force a WAL switch. The restore test in step 7 looks for the marker rows.
kubectl exec $PRIMARY -c postgres -- psql -U postgres -d postgres \
-c "CREATE TABLE IF NOT EXISTS backup_check (id serial PRIMARY KEY, note text, created_at timestamptz DEFAULT now());" \
-c "INSERT INTO backup_check (note) VALUES ('before-backup');" \
-c "SELECT pg_switch_wal();"
sleep 10Then run the archiver query again. These are the expected changes:
archived_countincreases.last_archived_walshows a newer file.failed_countdoes not increase. A small increase during the first minutes after a deployment is acceptable if it stops.
3. Replication
Read the replication status on the primary:
kubectl exec $PRIMARY -c postgres -- psql -U postgres -d postgres -c \
"SELECT application_name, state, sync_state, replay_lag FROM pg_stat_replication;"Healthy output has one row for each replica, all with the state streaming:
application_name | state | sync_state | replay_lag
------------------+-----------+------------+------------
pg-2 | streaming | async |
pg-3 | streaming | async |Then read the marker rows on a replica:
kubectl exec $REPLICA -c postgres -- psql -U postgres -d postgres -c "SELECT count(*) FROM backup_check;"The replica must return the same count as the primary. This result shows that the replication delivers data, not only a connection.
4. Plugin sidecar log
kubectl logs $PRIMARY -c plugin-barman-cloud --tail=50Healthy lines:
{"level":"info","ts":"...","msg":"Executing barman-cloud-wal-archive","logging_pod":"pg-1","walName":"/var/lib/postgresql/data/pgdata/pg_wal/000000010000000000000008",...}
{"level":"info","ts":"...","msg":"Archived WAL file","logging_pod":"pg-1"}
{"level":"info","ts":"...","msg":"Applying backup retention policy","logging_pod":"pg-1"}Problems and their causes:
| Message in the sidecar log | Cause and fix |
|---|---|
An error occurred (InvalidAccessKeyId) when calling the PutObject | The sidecar can reach S3, but the credentials are wrong. Correct the Secret that ObjectStore.spec.configuration.s3Credentials refers to. |
secrets "<name>" is forbidden, or other RBAC errors | The sidecar cannot read the Secret or the ObjectStore. The plugin manages this access. Contact k8s@cerit-sc.cz. |
MissingContentLength on UploadPart, or an x-amz-content-sha256 error | The S3 endpoint does not accept the checksums from boto3. Add the variables to instanceSidecarConfiguration.env in the ObjectStore. See Step 2 of the deploy guide. |
Many retries without Archived WAL file between them | The endpoint URL, the bucket name, or an S3 permission is wrong. |
By default, CloudNativePG makes base backups on a replica. The log lines of a base backup are in the sidecar of that replica.
5. Bucket
The steps above show that the process runs. Only the bucket shows that the files arrive.
The script below needs Python 3 and boto3 (pip install boto3). First, read the credentials from the Secret that the ObjectStore refers to. Use your Secret name and key names:
export AWS_ACCESS_KEY_ID=$(kubectl get secret pg-backup -o jsonpath='{.data.AWS_ACCESS_KEY_ID}' | base64 -d)
export AWS_SECRET_ACCESS_KEY=$(kubectl get secret pg-backup -o jsonpath='{.data.AWS_SECRET_ACCESS_KEY}' | base64 -d)Save this script as s3-list.py. It reads all objects of the Cluster, also when there are more than 1000:
#!/usr/bin/env python3
"""Show the base backups and the newest WAL files of a Cluster in S3.
Usage: python3 s3-list.py <bucket> <path>
<path> is the Cluster name, or <folder>/<cluster> if destinationPath has a folder.
"""
import os
import re
import sys
import boto3
# A WAL segment has a name of 24 hex digits, with an optional compression suffix.
WAL_FILE = re.compile(r"^[0-9A-F]{24}(\.[a-z0-9]+)?$")
bucket, path = sys.argv[1], sys.argv[2].strip("/")
s3 = boto3.client(
"s3",
endpoint_url="https://s3.a.cloud.e-infra.cz",
aws_access_key_id=os.environ["AWS_ACCESS_KEY_ID"],
aws_secret_access_key=os.environ["AWS_SECRET_ACCESS_KEY"],
)
backups, wals = set(), []
for page in s3.get_paginator("list_objects_v2").paginate(Bucket=bucket, Prefix=f"{path}/"):
for obj in page.get("Contents", []):
parts = obj["Key"][len(path) + 1:].split("/")
if parts[0] == "base" and len(parts) > 2:
backups.add(parts[1])
elif parts[0] == "wals" and WAL_FILE.match(parts[-1]):
wals.append(obj["Key"])
print("Base backups:")
for backup in sorted(backups):
print(f" {path}/base/{backup}/")
print(f"WAL files: {len(wals)}. Newest 5:")
for wal in sorted(wals, key=lambda key: key.rsplit("/", 1)[-1])[-5:]:
print(f" {wal}")Run the script:
python3 s3-list.py "$BUCKET" "$CLUSTER"If destinationPath has a folder, use "<folder>/$CLUSTER" as the second argument.
Healthy output:
Base backups:
pg/base/20261004T221041/
WAL files: 8. Newest 5:
pg/wals/0000000100000000/000000010000000000000004
pg/wals/0000000100000000/000000010000000000000005
pg/wals/0000000100000000/000000010000000000000006
pg/wals/0000000100000000/000000010000000000000007
pg/wals/0000000100000000/000000010000000000000008Run the script before and after a WAL switch (step 2). After the switch, the newest WAL file must have a higher number.
After a migration, the legacy base backups stay next to the new base backups. This is correct, because the plugin continues in the same path.
6. On-demand base backup
This backup does not change the scheduled backups.
apiVersion: postgresql.cnpg.io/v1
kind: Backup
metadata:
name: <cluster>-manual-1
spec:
cluster:
name: <cluster>
method: plugin
pluginConfiguration:
name: barman-cloud.cloudnative-pg.iokubectl apply -f backup.yaml
kubectl get backups.postgresql.cnpg.io $CLUSTER-manual-1
# NAME CLUSTER METHOD PHASE ERROR
# pg-manual-1 pg plugin completedThe phase changes from started to completed. Then run the S3 script again. A new folder <cluster>/base/<timestamp>/ must appear.
7. Restore test
A backup is good only if you can restore it. This test creates a temporary Cluster from the backups in the bucket. The test does not change the source Cluster.
-
Write a second marker row on the primary and force a WAL switch. This row is newer than the base backup. The restore gets it only if it replays the archived WAL files.
kubectl exec $PRIMARY -c postgres -- psql -U postgres -d postgres \ -c "INSERT INTO backup_check (note) VALUES ('after-backup');" \ -c "SELECT pg_switch_wal();" sleep 10 -
Run the archiver query from step 2. Make sure that
last_archived_walshows a newer file. -
Save this manifest as
restore-test.yaml. Use the same image as the source Cluster. SetserverNameto the name of the source Cluster. Setdatabaseandownerto the values of the source Cluster.restore-test.yaml apiVersion: postgresql.cnpg.io/v1 kind: Cluster metadata: name: <cluster>-restore-test spec: instances: 1 # Use the same image as the source Cluster. imageName: ghcr.io/cloudnative-pg/postgresql:15-standard-trixie # Do not create the read-only Services. # Remove this block only if you need read-only connections to the replicas. managed: services: disabledDefaultServices: ["r", "ro"] storage: size: 5Gi # at least the size of the source volume storageClass: zfs-csi bootstrap: recovery: source: origin database: <db> # database of the source Cluster owner: <owner> # owner of that database externalClusters: - name: origin plugin: name: barman-cloud.cloudnative-pg.io parameters: barmanObjectName: pg-backup-store serverName: <cluster> # No "plugins" section: the test Cluster must not archive WAL files. -
Apply the manifest. Then wait until the test Cluster is ready. The time depends on the size of the database.
kubectl apply -f restore-test.yaml kubectl wait cluster.postgresql.cnpg.io/$CLUSTER-restore-test --for=condition=Ready --timeout=1800s -
Read the marker rows in the restored database:
kubectl exec $CLUSTER-restore-test-1 -c postgres -- psql -U postgres -d postgres -c \ "SELECT note, created_at FROM backup_check ORDER BY id;"The result must contain
before-backupandafter-backup. The rowbefore-backupcomes from the base backup. The rowafter-backupcomes from the WAL archive. -
Delete the test Cluster. The operator also deletes its volume.
kubectl delete cluster.postgresql.cnpg.io $CLUSTER-restore-test
Notes for the restore test:
- The test Cluster uses resources from your namespace quota. Make sure that the quota has space for one more instance.
- Do not add a
pluginssection to the test Cluster. If you add it, the test Cluster writes its own WAL archive to the bucket. The next restore test then fails the empty-archive check. - If a NetworkPolicy in your namespace denies traffic by default, allow the same traffic for the test Cluster. See Step 5 of the deploy guide.
- To restore to a point in time, add
recoveryTargetunderbootstrap.recovery, for examplerecoveryTarget: { targetTime: "2026-10-04 22:15:00+00" }.
Periodic checks
Add these values to your monitoring:
barman_cloud_cloudnative_pg_io_last_failed_backup_timestamp: must not change.barman_cloud_cloudnative_pg_io_last_available_backup_timestamp: must follow the schedule.barman_cloud_cloudnative_pg_io_first_recoverability_point: must move forward when the retention policy deletes old backups.cnpg_pg_stat_archiver_failed_count: must not increase.cnpg_pg_stat_archiver_seconds_since_last_archival: must stay low on a database with write activity. A high value can mean that WAL files collect in the volume.
The first three metrics come from the plugin. The legacy in-tree backup uses other names: cnpg_collector_last_failed_backup_timestamp, cnpg_collector_last_available_backup_timestamp, and cnpg_collector_first_recoverability_point.
Also repeat the restore test at regular intervals, for example one time each month.
