Backups & Restore
Esta página aún no está disponible en tu idioma.
GeoLens ships an automated backup container that runs on a configurable cron schedule, retains daily and weekly snapshots locally, and optionally replicates to S3-compatible storage. There is no /admin/backups UI page. Backups are operated entirely from the Docker Compose CLI and via environment variables.
Replace https://geolens.example.com with your GeoLens instance’s URL in every example below.
Automated backups are on by default
Section titled “Automated backups are on by default”The backup container starts with the default Docker Compose stack. A plain docker compose up -d runs scheduled pg_dump backups, keeps 7 daily snapshots and 4 weekly snapshots, and optionally uploads to S3. The service uses the same database credentials as the API, so no additional configuration is required for the local-only path. Prebuilt installs pull the published multi-arch ghcr.io/geolens-io/geolens-backup image at the pinned release version; source installs build the same stage from the root Dockerfile.
To verify the backup container is running:
docker compose ps backupThe healthcheck measures freshness, not liveness: the container is healthy only while the last fully successful backup cycle (dump, verification, staging archive, and S3 upload when enabled) is younger than BACKUP_MAX_AGE_MINUTES (default 1560 minutes, i.e. 26 hours: the daily schedule plus slack). A container whose backups quietly stop succeeding turns unhealthy roughly one missed cycle later. If you run a less-frequent BACKUP_SCHEDULE, raise BACKUP_MAX_AGE_MINUTES to about 1.5x the interval. On a fresh install the service shows starting until the initial on-startup backup succeeds.
The backup container restarts automatically; if it exits non-zero (e.g., disk full, connectivity issue), Docker restarts it and the next scheduled run resumes normally.
Configuration
Section titled “Configuration”| Variable | Default | Purpose |
|----------|---------|---------|
| BACKUP_SCHEDULE | 0 2 * * * | Cron expression for backup execution |
| BACKUP_RETENTION_DAILY | 7 | Number of daily backups to keep locally |
| BACKUP_RETENTION_WEEKLY | 4 | Number of weekly (Sunday) backups to keep locally |
| BACKUP_S3_ENABLED | false | Upload backups to S3 in addition to local storage |
| S3_ENDPOINT, S3_BUCKET, S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY, S3_REGION | (none) | S3 credentials (shared with the storage provider) |
| S3_ADDRESSING_STYLE | auto | S3 addressing style (auto, path, virtual); use path for MinIO |
| BACKUP_MAX_AGE_MINUTES | 1560 | Healthcheck freshness threshold: the container reports unhealthy when the last fully successful backup cycle is older than this many minutes |
| BACKUP_MEM_LIMIT | 512m | Container memory cap for the backup service (the nightly pg_dump + tar peaks around 300 MB) |
Set these in .env before starting the stack. Changes require a container restart:
docker compose restart backupThe backup service reads the same S3_* credentials as the storage provider. There are no dedicated BACKUP_S3_* bucket/credential overrides. BACKUP_S3_ENABLED is the only backup-specific S3 knob; off-site upload reuses S3_ENDPOINT, S3_BUCKET, S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY, and S3_REGION. To back up to a separate bucket, point those shared S3_* values at it. The uploader follows whatever scheme S3_ENDPOINT carries, so use an http:// endpoint for a plain-HTTP MinIO; S3_ALLOW_HTTP only affects the app’s own object-storage client, not the backup uploader. See Configuration Reference.
Backup destinations
Section titled “Backup destinations”Backups are written to:
/backups/daily/<db>_<timestamp>.dump: every scheduled run/backups/weekly/<db>_<timestamp>.dump: Sundays only/backups/{daily,weekly}/staging-<timestamp>.tar.gz: the localupload_stagingvolume (rasters/COGs/source uploads), tarred alongside each dump with the same timestamp so a restore can recover a working instance rather than an empty-object catalog/backups/{daily,weekly}/globals-<timestamp>.sql: the cluster globals (pg_dumpall --globals-only) — roles, their passwords, and cluster-wide grants — written next to the dump with the same timestamps3://<bucket>/backups/{daily,weekly}/...: the dump, the staging archive, and the globals dump, whenBACKUP_S3_ENABLED=true
The <db> segment is the database name (default geolens). The <timestamp> segment is YYYYMMDD_HHMMSS in UTC. Dumps are PostgreSQL custom-format (pg_dump -Fc --no-owner --no-acl). Restore with pg_restore rather than psql.
Local backups live inside the backup_data Docker volume. Inspect them by execing into the backup container:
docker compose exec backup ls -lh /backups/daily/docker compose exec backup ls -lh /backups/weekly/To copy a specific backup off the host, take the whole set — all three files share one timestamp:
# umask 077: the globals file holds role password verifiers, so the copies# must not be readable by other users on the host either.(umask 077 docker compose cp backup:/backups/daily/geolens_20260101_020000.dump ./ docker compose cp backup:/backups/daily/staging-20260101_020000.tar.gz ./ docker compose cp backup:/backups/daily/globals-20260101_020000.sql ./)The staging archive is absent on a fresh install with no uploaded datasets; the globals dump is not optional, and a cycle that fails to produce it is reported as a failed cycle.
Retention policy
Section titled “Retention policy”The backup container enforces retention on every run. After producing a new backup:
- Daily backups older than
BACKUP_RETENTION_DAILYruns are deleted from/backups/daily/. - Weekly backups older than
BACKUP_RETENTION_WEEKLYruns are deleted from/backups/weekly/. - S3 retention is not enforced by GeoLens. Configure S3 lifecycle rules on the bucket itself for off-site retention. The default 30-day Glacier transition + 365-day expiry is a common starting point for compliance-driven deployments.
Manual deletion is safe: the container scans the directory on the next run and updates retention based on what’s present. Out-of-band file deletion does not corrupt the backup state.
Both retention values must be plain integers of at least 1. The container refuses to start otherwise (the log shows must be >= 1), because a retention of 0 would delete each backup the moment it is written.
Manual pg_dump (alternate path)
Section titled “Manual pg_dump (alternate path)”For one-off backups outside the scheduled cron, run pg_dump directly against the database container:
# Create a custom-format dump (preferred for pg_restore). Mirror the# automated backup's flags so it restores cleanly via scripts/restore.shdocker compose exec db pg_dump -U geolens -d geolens -Fc --no-owner --no-acl -f /tmp/geolens_backup.dumpdocker compose cp db:/tmp/geolens_backup.dump ./geolens_backup.dumpFor a plain SQL backup (useful for grep/inspection):
docker compose exec db pg_dump -U geolens -d geolens > geolens_backup.sqlFor volume-level snapshots (database files plus WAL):
# Stop services first to ensure a consistent snapshotdocker compose down
# Backup the pgdata volumedocker run --rm -v geolens_pgdata:/data -v $(pwd):/backup alpine \ tar czf /backup/pgdata_backup.tar.gz -C /data .
# Restartdocker compose up -dVolume snapshots are faster to restore than pg_restore for full-instance recovery but cannot be restored to a different PostgreSQL version. Use the pg_dump format for long-term archival and cross-version portability.
Restoring from a backup
Section titled “Restoring from a backup”The canonical restore path is scripts/restore.sh. It validates the dump with pg_restore -f /dev/null before touching the database (reading every data block; a merely truncated dump still passes --list), reports any sibling globals-<timestamp>.sql and names the roles it defines that the target cluster is missing, pre-creates the required extensions/schemas/reader role, stops api and worker to prevent write conflicts, runs the restore, then restarts the services via an EXIT trap (even on failure):
# Restores into the running db container from a custom-format dump./scripts/restore.sh /path/to/geolens_20260101_020000.dumpInternally it runs:
pg_restore -U "$POSTGRES_USER" -d "$POSTGRES_DB" --clean --if-exists --no-owner < "$BACKUP_FILE"The --clean --if-exists flags drop existing objects first and tolerate objects that are absent on a fresh database (those produce warnings, not errors; the script treats a warning-only nonzero exit as success). --no-owner skips ownership commands so the dump restores cleanly under the geolens role regardless of who created it.
To restore the DB by hand (e.g. without the wrapper), match those flags:
# Stop services to prevent writesdocker compose stop api worker
# Copy the backup into the database containerdocker compose cp ./backup.dump db:/tmp/backup.dump
# Restore (mirror restore.sh's flags)docker compose exec db pg_restore -U geolens -d geolens \ --clean --if-exists --no-owner /tmp/backup.dump
# Restartdocker compose start api workerFor partial restores (single table), add --table <name>:
docker compose exec db pg_restore -U geolens -d geolens \ --clean --if-exists --no-owner --table catalog.datasets /tmp/backup.dumpTo restore a volume-level snapshot:
docker compose downdocker run --rm -v geolens_pgdata:/data -v $(pwd):/backup alpine \ sh -c "rm -rf /data/* && tar xzf /backup/pgdata_backup.tar.gz -C /data"docker compose up -dRestore validation
Section titled “Restore validation”restore.sh ends with a row-count sanity check (catalog.records and catalog.datasets), but beyond that restore validation is operator-driven: there is no automated post-restore checksum or integrity verification. Recommended practice:
-
Restore the most recent daily backup into a staging instance monthly.
-
Run
curl https://staging.geolens.example.com/api/healthand confirmdatabase.status: ok(see Infrastructure & Monitoring for the health probe). Use/api/health, not a bare/health: through the bundled Nginx the latter returns the SPA’s HTML with status 200 and no provider JSON. -
Spot-check 2 to 3 dataset rows from the catalog API:
Terminal window curl https://staging.geolens.example.com/api/datasets/ \-H "Authorization: Bearer $TOKEN" | jq '.datasets[:3]' -
Open a representative dataset in the staging UI and confirm map preview, exports, and metadata render correctly.
If staging restore reveals corruption, the most recent daily backup may be unrecoverable. Promote the most recent weekly backup to production-restore status and investigate the failed daily as a separate incident.
See also
Section titled “See also”- Settings reference: Storage tab covers the
S3_*credentials the backup service reuses whenBACKUP_S3_ENABLED=true - Self-hosted provider notes: backups and S3 off-site replication with managed providers
- Configuration reference: full
BACKUP_*env-var details - RUNBOOK.md: the operator runbook in the repository — fresh-cluster and managed-Postgres restore, data-loss incident response, schema rollback, and PostgreSQL major-version upgrades
- Upgrade guide: version-specific operator notes, including the 1.7.0 release that introduced the
globals-*.sqlartifact