Runbooks

Server access

  • Both servers: SSH as root with the admin key (ivydigital-admin-ed25519), kept on the operator PC.
  • dev server: port 2222 (port 22 blocked) · ssdserver: port 22.
  • Windows OpenSSH works on both (plink only if needed on dev).

DNS & certificates

  • Cloudflare dashboard for the ivydigitals.com zone; API token lives in /root/.secrets/cloudflare.ini (has DNS edit; cache-purge is manual).
  • New cert: certbot certonly --dns-cloudflare --dns-cloudflare-credentials /root/.secrets/cloudflare.ini -d <host> --non-interactive --agree-tos
  • Renewal is automatic (certbot timer). Certs expire ~90 days after issue.

Backups & restore

  • Backups exist from every migration: DB dumps + code tars with SHA256 checksums in /root/migration-* on both servers.
  • Databases: Postgres via pg_dump, MySQL via mysqldump, SQLite via sqlite3 ".backup" (safe with WAL).
  • Restore = recreate DB → import dump → restore code tar → restart services. The per-app docs pages list exact DB names.
  • Recommended next step: schedule nightly backups of all databases + /root/migration dirs to off-server storage.

Monitoring

  • ssdserver runs the monitoring stack: Prometheus, Grafana, Loki, cAdvisor, node-exporter (ports 9000-9201 range are used — e.g. 9100 is node-exporter).
  • Grafana dashboards for both servers were set up in earlier infra work.

Incident response

  1. Check the app in the Release Center (health dots).
  2. If a release caused it: Rollback in the Release Center.
  3. If the app is down: check PM2 (pm2 list / pm2 logs <name>) or Docker (docker ps, docker logs <name>) on the relevant server.
  4. Web 502/500: check nginx error log /var/log/nginx/error.log and the upstream port.
  5. DB issues: verify the DB container is up and the app's connection string matches (see per-app pages).

Staging refresh

  • Staging DBs are point-in-time snapshots from the migration. To refresh staging from prod: dump prod DB → import into staging DB → re-apply staging URL overrides → keep dangerous crons disabled.
  • Dangerous crons stay OFF on staging (lumen autopilot, ivyos agent fetches) — enabling them would perform real actions.

IVY Digital · docs.ivydigitals.com · generated 2026-10-01 · Release Center · Gitea