# Production deployment The production stack contains Nginx, the React web build, the Fastify API, a BullMQ worker, PostgreSQL, Redis, MinIO, Prometheus, and Alertmanager. API and storage services are not published directly. Nginx routes the application host to the SPA/API and the media host to MinIO. ## Prerequisites - Docker Engine with Compose v2 - Two DNS records pointing to the deployment host, such as `studio.example.com` and `media.example.com` - An HTTPS reverse proxy or load balancer in front of port `8080` - OpenAI and Replicate credentials for real generation jobs - A Stripe account, recurring Price IDs, and webhook signing secret when subscription billing is enabled - An HTTPS operator webhook that accepts Alertmanager webhook payloads - An encrypted off-host destination for daily backups Both public hosts must be forwarded to the same Nginx port with the original `Host` header preserved. TLS is required because production refresh cookies are secure. `S3_PUBLIC_ENDPOINT` must exactly match the public media origin; providers and browsers cannot use the private `http://minio:9000` address. Production registration requires verified email by default. Configure `RESEND_API_KEY` and a verified `EMAIL_FROM` sender before startup; password-reset, verification, and workspace-invitation links use the first origin in `WEB_ORIGIN` as their public base URL. Workspace invitations remain valid when email delivery fails and the API returns a one-time manual link, but operators should treat a `FAILED` delivery result as an email-provider incident rather than asking the inviter to recreate the workspace. Subscription billing remains read-only when Stripe is not configured. To enable checkout, set `STRIPE_SECRET_KEY`, `STRIPE_WEBHOOK_SECRET`, `STRIPE_PRO_PRICE_ID`, and `STRIPE_STUDIO_PRICE_ID`, then register `https://studio.example.com/api/v1/billing/webhooks/stripe` for `checkout.session.completed` and `customer.subscription.created`, `customer.subscription.updated`, and `customer.subscription.deleted`. Configure the Stripe Customer Portal before exposing subscription management; `STRIPE_PORTAL_CONFIGURATION_ID` is optional when the account default configuration is used. Before enabling billing, run the read-only Stripe resource smoke. It never creates a customer, session, charge, subscription, or event. Use `--require-live` for the production account: ```bash FRAMEFLOW_STRIPE_APP_ORIGIN=https://studio.example.com \ STRIPE_SECRET_KEY='use-the-secret-store-value' \ STRIPE_WEBHOOK_SECRET='use-the-secret-store-value' \ STRIPE_PRO_PRICE_ID=price_replace_with_pro \ STRIPE_STUDIO_PRICE_ID=price_replace_with_studio \ STRIPE_PORTAL_CONFIGURATION_ID=bpc_replace_when_not_using_default \ npm run smoke:stripe -- --require-live --json ``` The smoke retrieves the current account, expands both configured Products from their recurring Prices, validates active licensed recurring billing and live/test mode consistency, finds the explicit or default active Customer Portal configuration, and locates the exact enabled FrameFlow webhook with all four required events. The secret key and webhook signing secret are read only from the environment and are never printed. Stripe does not expose an endpoint's signing secret through its API, so this check can prove that `STRIPE_WEBHOOK_SECRET` is configured but cannot prove it belongs to the discovered endpoint. Complete one signed test subscription event through the FrameFlow webhook and confirm the workspace plan/credit update before accepting billing. ## Publishing connector Direct delivery to Douyin, Kuaishou, Bilibili, and Xiaohongshu is intentionally isolated behind an operator-owned connector. The connector holds platform application credentials and translates the stable FrameFlow contract into each platform's OAuth, upload, polling, and token-refresh APIs. Without it, download delivery remains available and the publishing UI reports every external channel as not configured. Configure these values together: - `API_PUBLIC_ORIGIN`: the public HTTPS FrameFlow API origin used to build OAuth and status callback URLs. - `CREDENTIAL_ENCRYPTION_KEY`: a production-only secret of at least 32 characters used to encrypt channel access tokens, refresh tokens, and OAuth PKCE verifiers at rest. - `PUBLISHING_CONNECTOR_URL`: the connector's public HTTPS origin. - `PUBLISHING_CONNECTOR_SECRET`: a shared secret of at least 24 characters. FrameFlow sends it as `Authorization: Bearer ` and also uses it to verify connector callbacks. - `PUBLISHING_CONNECTOR_TIMEOUT_MS`: timeout for each connector request, from 1,000 to 120,000 milliseconds. Every connector request is JSON and includes `platform`, one of `DOUYIN`, `KUAISHOU`, `BILIBILI`, or `XIAOHONGSHU`. The connector must expose: - `POST /v1/oauth/authorize`: accepts `{ platform, state, callbackUrl }`; returns `{ authorizationUrl, codeVerifier? }`. The platform must return the supplied `state` to `callbackUrl`. - `POST /v1/oauth/exchange`: accepts `{ platform, code, callbackUrl, codeVerifier? }`; returns `{ externalAccountId, displayName, avatarUrl?, accessToken, refreshToken?, tokenExpiresAt?, scopes?, metadata? }`. - `POST /v1/oauth/revoke`: accepts `{ platform, externalAccountId, accessToken, refreshToken? }`; returns `{ revoked: boolean }`. - `POST /v1/publish`: accepts `{ platform, deliveryId, sourceUrl, callbackUrl, account, destination }`; returns `{ status, externalId, publishedUrl?, credentials? }`. `status` is `PUBLISHING` for asynchronous processing or `PUBLISHED` only when the platform has confirmed publication. `sourceUrl` is a short-lived signed media URL and must not be persisted or logged. For asynchronous work, POST status events to the supplied FrameFlow `callbackUrl` with this body: ```json { "eventId": "provider-unique-event-id", "platform": "DOUYIN", "deliveryId": "00000000-0000-0000-0000-000000000000", "externalId": "platform-video-id", "status": "PUBLISHED", "occurredAt": "2026-07-31T12:00:00.000Z", "publishedUrl": "https://platform.example/video/123" } ``` `status` may be `PUBLISHING`, `PUBLISHED`, or `FAILED`; include `errorMessage` for failures. Sign the exact raw JSON bytes with HMAC-SHA256 using `PUBLISHING_CONNECTOR_SECRET`, and send the lowercase hexadecimal digest as `x-frameflow-signature: sha256=`. FrameFlow deduplicates events by platform and event ID, rejects stale or mismatched events, and never allows a terminal delivery to regress. Before enabling a channel in production, bind and revoke a dedicated test account, publish an approved test render, wait for a signed terminal callback, open the returned platform URL, verify title/description/visibility/tags/schedule behavior, and confirm that refreshed credentials remain usable. Configuration-only readiness or a successful queue submission is not acceptance evidence. ## Continuous integration `.github/workflows/ci.yml` provides two required pre-deployment checks. The verification job starts isolated PostgreSQL, Redis, and MinIO services, applies every migration, runs all web and server tests, creates a real backup of those services, restores it through the isolated rehearsal path, builds both applications, and audits production dependencies. The container job validates the production Compose model, builds the server and web images, confirms a synthetic safe production configuration is accepted, and confirms `.env.production.example` is rejected because it still contains placeholders. The synthetic values exist only inside the CI runner and are not deployment secrets. CI proves the repository is internally buildable and the safety gate works; it does not prove external provider credentials, public TLS/media routing, email delivery, billing webhooks, alert delivery, or cloud generation jobs. Complete the live checks below for every release candidate deployed to a real environment. ## Production smoke test Use a dedicated, verified account with the `OWNER` or `ADMIN` role. The smoke runner accepts its password only through the environment, requires HTTPS outside localhost, and does not disable certificate validation. Its default mode never starts OpenAI or Replicate jobs. It performs these checks through the public application origin: - API liveness plus PostgreSQL, Redis, object-storage, and worker readiness - authenticated service readiness including FFmpeg, AI configuration, and any persisted provider-verification state without starting a new provider request - billing/usage read-model access - project, episode, immutable script, asset, and shot-version persistence - multipart object upload plus signed-download SHA-256 integrity - a real Redis/BullMQ worker health job - a 640x360 FFmpeg render, immutable source manifest, SRT/VTT sidecars, and signed video download - render review approval, audit-log persistence, download delivery, redirect handling, and delivered-video integrity - durable render/delivery notifications, user scoping, and persisted read state ```bash FRAMEFLOW_SMOKE_BASE_URL=https://studio.example.com \ FRAMEFLOW_SMOKE_EMAIL=production-smoke@example.com \ FRAMEFLOW_SMOKE_PASSWORD='use-the-secret-store-value' \ FRAMEFLOW_SMOKE_WORKSPACE_ID=00000000-0000-0000-0000-000000000000 \ npm run smoke:production -- --json ``` The workspace ID is optional when the account belongs to exactly one workspace. The default per-job timeout is 180 seconds and can be changed with `--timeout-seconds`. The test creates a uniquely marked project and archives it in a `finally` cleanup path on success or failure; use `--keep-project` only for diagnosis. Archiving preserves its database history and private media objects, so apply an operator-reviewed retention policy to old `[SMOKE]` projects rather than granting the test runner destructive storage permissions. For local-only verification, explicitly allow an HTTP loopback origin: ```bash FRAMEFLOW_SMOKE_BASE_URL=http://127.0.0.1:8787 \ FRAMEFLOW_SMOKE_ALLOW_HTTP=true \ FRAMEFLOW_SMOKE_EMAIL=studio@frameflow.local \ FRAMEFLOW_SMOKE_PASSWORD='FrameFlow2026!' \ npm run smoke:production -- --timeout-seconds 120 ``` A passing default report proves the local production path, storage, queue, render, review, and delivery adapters worked for that deployment. It never contacts an AI provider. AI readiness is configuration-only unless an owner or administrator previously completed the provider connection check; persisted connection status still does not replace the billable generation acceptance run. To validate the live provider path, add `--include-ai` and the exact billable-charge acknowledgement. Both are required; setting the environment acknowledgement alone never enables provider calls. Run this only in a workspace where test media and provider charges are acceptable: ```bash FRAMEFLOW_SMOKE_BASE_URL=https://studio.example.com \ FRAMEFLOW_SMOKE_EMAIL=production-smoke@example.com \ FRAMEFLOW_SMOKE_PASSWORD='use-the-secret-store-value' \ FRAMEFLOW_SMOKE_WORKSPACE_ID=00000000-0000-0000-0000-000000000000 \ FRAMEFLOW_SMOKE_BILLABLE_AI_ACK=I_ACCEPT_BILLABLE_AI_CHARGES \ FRAMEFLOW_SMOKE_AI_VIDEO_SECONDS=5 \ npm run smoke:production -- --include-ai --json ``` The live preflight runs before project creation and requires all seven workflows (`script`, `storyboard`, `image-openai`, `image-replicate`, `tts`, `video`, and `lipsync`) to be `CONFIGURED`, plus at least 700 remaining workspace credits. It then creates one OpenAI character image, one Replicate consistency variant, an OpenAI script and storyboard, OpenAI speech, Replicate video and lipsync media, and an FFmpeg render. Every provider job uses `maxAttempts: 1`; the report verifies stored media, provider/model attribution, the final render, and one usage-ledger entry per provider job. The uniquely marked smoke project is archived in the same cleanup path used by default mode unless `--keep-project` is supplied. `FRAMEFLOW_SMOKE_AI_VIDEO_SECONDS` accepts 1 to 30 seconds and defaults to 5. Some configured Replicate model versions require additional input fields. Supply only fields documented by those exact model versions as JSON objects through `FRAMEFLOW_SMOKE_VIDEO_PARAMS_JSON` and `FRAMEFLOW_SMOKE_LIPSYNC_PARAMS_JSON`; the runner rejects malformed JSON and arrays before login. Leave them unset when the selected versions need no extra parameters. Do not put provider tokens or other secrets in these JSON values or command-line arguments. Example Caddy routing in front of the stack: ```caddyfile studio.example.com { reverse_proxy 127.0.0.1:8080 } media.example.com { reverse_proxy 127.0.0.1:8080 } ``` The web client keeps one authenticated server-sent event stream open at `/api/v1/workspaces/:workspaceId/events`. The supplied API response disables proxy buffering and sends a heartbeat every 20 seconds. Preserve the `X-Accel-Buffering: no` response header, keep the upstream read timeout above 60 seconds, and do not enable response caching or compression that buffers `text/event-stream`. Notification history and read state are persisted in PostgreSQL, so reconnecting clients recover missed events through the regular notification endpoint; the existing task poller remains a degraded-mode fallback when streaming is unavailable. ## First deployment 1. Create the production environment file and replace every placeholder secret. ```bash cp .env.production.example .env.production openssl rand -hex 32 ``` 2. Create the host directory shared by the backup script and Node Exporter. If `FRAMEFLOW_BACKUP_METRICS_DIR` is changed in `.env.production`, use the same path here and in the backup scheduler environment. ```bash sudo install -d -m 0755 -o "$(id -un)" -g "$(id -gn)" /var/lib/frameflow/metrics ``` 3. Validate interpolation, build the lightweight configuration gate, and run it before starting any stateful application service. ```bash docker compose --env-file .env.production -f docker-compose.prod.yml config --quiet docker compose --env-file .env.production -f docker-compose.prod.yml build config-check docker compose --env-file .env.production -f docker-compose.prod.yml run --rm --no-deps config-check ``` The gate rejects development defaults, placeholder credentials, identical JWT signing secrets, localhost or non-HTTPS public endpoints, unauthenticated database/Redis URLs, Replicate model versions without a token, and partially configured Stripe billing. It prints only a sanitized capability summary and never logs credential values. The migration service also depends on this check, so an unsafe configuration cannot migrate or start the API through the supplied Compose stack. 4. Build and start the stack. Database migrations and bucket initialization run as one-shot dependencies before the API and worker start. ```bash docker compose --env-file .env.production -f docker-compose.prod.yml up -d --build docker compose --env-file .env.production -f docker-compose.prod.yml ps ``` 5. Verify public and internal health. ```bash curl -fsS https://studio.example.com/health/live curl -fsS https://studio.example.com/health/ready docker compose --env-file .env.production -f docker-compose.prod.yml exec api node -e "fetch('http://127.0.0.1:8787/health/ready').then(r=>r.text()).then(console.log)" docker compose --env-file .env.production -f docker-compose.prod.yml exec worker node -e "fetch('http://127.0.0.1:9092/health/ready').then(r=>r.text()).then(console.log)" ``` API readiness includes PostgreSQL, Redis, object storage, and the Redis-backed worker heartbeat. A running API returns `503 not_ready` when the generation worker is absent or stale, so the web service is not promoted while production jobs cannot run. 6. Create the first workspace owner through the registration API, then use that account on the login screen. ```bash curl -fsS https://studio.example.com/api/v1/auth/register \ -H 'content-type: application/json' \ --data '{"email":"owner@example.com","password":"replace-with-a-long-password","displayName":"Studio Owner","workspaceName":"My Studio"}' ``` Public registration should be disabled or protected at the edge after account provisioning when operating an invite-only deployment. ## Media and model checks MinIO CORS is restricted to `WEB_ORIGIN` for browser requests. CORS does not make objects public; every object remains private and requires a short-lived signature. Before accepting production work, verify one job of every configured provider type: - script generation and storyboard generation - character image generation using reference asset versions - speech generation - image-to-video generation - lipsync generation - episode render, immutable source-manifest capture, SRT/VTT subtitle download, and approved video delivery Workspace owners and administrators can run a non-generating connection check from the AI Services page. OpenAI verification retrieves every distinct configured text, image, and speech model through the Models API. Replicate verification reads the authenticated account endpoint, so it proves the token is accepted but does not prove that a configured model version can complete a prediction. Results are stored per workspace, audited, and automatically ignored after a credential or model configuration change. Provider response bodies and credentials are never stored. Complete the live billable workflow below before treating generation as accepted. Replicate must be able to fetch the public media origin over HTTPS. Provider output downloads are accepted only from public HTTPS addresses; every redirect is revalidated, streamed data is capped by `MAX_UPLOAD_BYTES`, and `PROVIDER_DOWNLOAD_TIMEOUT_MS` plus `PROVIDER_DOWNLOAD_MAX_REDIRECTS` bound stalled or redirecting responses. Do not mark provider validation complete when only queue creation succeeds; inspect the resulting media and usage ledger entry. `REPLICATE_PREDICTION_DEADLINE_SECONDS` controls both the local polling deadline and Replicate's `Cancel-After` deadline. It accepts 5 to 86,400 seconds and defaults to 900. Set it high enough for the selected model versions, but keep it bounded so a worker timeout does not leave a remote prediction running and accruing charges. ## Monitoring Prometheus listens on `127.0.0.1:9090` by default and scrapes API, worker, host, and MinIO metrics over the private Compose network. MinIO's metrics endpoint is unauthenticated only inside that network; Nginx returns `404` for both v2 and v3 metrics paths on the public media host. Alertmanager listens on `127.0.0.1:9093` and sends grouped firing and resolved notifications to `ALERTMANAGER_WEBHOOK_URL`. The URL is rendered into a temporary in-container config and must use HTTPS. Keep `METRICS_TOKEN` empty with the supplied scrape configuration. If metrics are exposed outside the Compose network, configure a token and update Prometheus authorization at the same time. Useful checks: ```bash curl -fsS http://127.0.0.1:9090/-/ready curl -fsS http://127.0.0.1:9093/-/ready curl -fsS 'http://127.0.0.1:9090/api/v1/targets' curl -fsS 'http://127.0.0.1:9090/api/v1/rules' curl -fsS 'http://127.0.0.1:9090/api/v1/alerts' curl -fsS 'http://127.0.0.1:9090/api/v1/alertmanagers' docker compose --env-file .env.production -f docker-compose.prod.yml logs --since=30m api worker node-exporter prometheus alertmanager ``` The supplied `deploy/alerts.yml` rules cover API/worker/exporter disappearance, stale worker heartbeat, dependency readiness, host and PostgreSQL-volume disk pressure, MinIO usable capacity, backup age, restore-rehearsal age, API error rate and latency, generation failure rate, and queue delay. Prometheus evaluates these rules and routes alerts to the bundled Alertmanager. The exact `FrameFlowDeliveryTest` alert uses a dedicated one-second group wait and five-second group interval so production acceptance does not change the timing of operational alerts. Run the controlled delivery smoke before launch. The exact acknowledgement is required because this command sends real firing and resolved notifications to the configured operator webhook: ```bash FRAMEFLOW_ALERTMANAGER_URL=http://127.0.0.1:9093 \ FRAMEFLOW_ALERT_TEST_ACK=I_ACCEPT_TEST_ALERT_NOTIFICATIONS \ npm run smoke:alertmanager -- --json ``` The production Alertmanager enables its supported `receiver-name-in-metrics` feature so the command can scope notification counters to `frameflow-operator-webhook`. The command records those counters, posts a uniquely labeled firing alert, waits for the alert to become active and for a successful webhook attempt, resolves the same alert, then waits for both the inactive state and a second successful webhook attempt. It fails when Alertmanager reports a failed notification and always attempts to resolve a firing test alert after an intermediate error. A passing result proves that the webhook returned success for both payloads. Use the returned `testId` to confirm the receiver performed its own downstream processing; Alertmanager cannot prove work performed after the receiver acknowledged the HTTP request. ## Backups The backup script creates a PostgreSQL custom-format dump plus Redis and MinIO volume archives, then writes checksums. Before publishing success metrics it runs the read-only backup verifier, which checks the required file set, manifest shape, exact checksum inventory, PostgreSQL custom-format signature, readable gzip archives, and unsafe absolute or parent-directory archive paths. Only after every check succeeds does it atomically publish backup timestamp, duration, and size metrics for Node Exporter. Run it from the repository root or a daily scheduler with an explicit off-host destination. A failed run leaves the previous success timestamp unchanged so the stale-backup alert can fire. ```bash FRAMEFLOW_ENV_FILE=/srv/frameflow/.env.production \ FRAMEFLOW_BACKUP_METRICS_DIR=/var/lib/frameflow/metrics \ /srv/frameflow/scripts/backup.sh /mnt/encrypted-backups/frameflow ``` After the first successful backup, verify the metric through Prometheus: ```bash curl -fsS --get http://127.0.0.1:9090/api/v1/query \ --data-urlencode 'query=frameflow_backup_last_success_timestamp_seconds' ``` Copy completed backup directories off the application host. A local Docker volume is not a disaster-recovery backup. Test checksum validation and restore into an isolated Compose project on a regular schedule. ### Scheduled backup and restore rehearsal The supplied systemd units assume a system Docker daemon, an application checkout at `/srv/frameflow`, a `frameflow` service account with supplementary membership in the `docker` group, an encrypted mount at `/mnt/encrypted-backups`, and the standard metrics directory. Change every matching path in both service units when the host layout differs; changing only the environment file is insufficient because the systemd filesystem allowlist and mount assertion are intentionally explicit. Create the service account according to the host's account-management policy, mount the encrypted off-host filesystem through `/etc/fstab` or a dedicated mount unit, then install the configuration: ```bash sudo install -d -m 0750 -o root -g frameflow /etc/frameflow sudo install -d -m 0750 -o frameflow -g frameflow /var/lib/frameflow/metrics sudo install -d -m 0700 -o frameflow -g frameflow /mnt/encrypted-backups/frameflow sudo chown root:frameflow /srv/frameflow/.env.production sudo chmod 0640 /srv/frameflow/.env.production sudo install -m 0640 -o root -g frameflow \ deploy/systemd/frameflow-backup.env.example /etc/frameflow/backup.env sudo install -m 0644 deploy/systemd/*.service deploy/systemd/*.timer /etc/systemd/system/ sudoedit /etc/frameflow/backup.env sudo systemd-analyze verify \ /etc/systemd/system/frameflow-backup.service \ /etc/systemd/system/frameflow-backup.timer \ /etc/systemd/system/frameflow-restore-rehearsal.service \ /etc/systemd/system/frameflow-restore-rehearsal.timer sudo systemctl daemon-reload sudo systemctl enable --now frameflow-backup.timer frameflow-restore-rehearsal.timer ``` `scheduled-maintenance.sh` resolves both configured paths, rejects relative paths, root, whitespace, symbolic links, destinations outside the mount, and inactive mount points. This prevents an unavailable encrypted mount from silently redirecting backups onto the application host. The monthly restore unit requires the daily backup service and waits for a fresh successful backup before selecting the newest timestamped directory. It verifies that directory once before the restore script verifies it again immediately before any Docker call. Run both units manually during first deployment and inspect their logs and timers: ```bash sudo systemctl start frameflow-backup.service sudo systemctl start frameflow-restore-rehearsal.service sudo systemctl status frameflow-backup.service frameflow-restore-rehearsal.service sudo systemctl list-timers frameflow-backup.timer frameflow-restore-rehearsal.timer journalctl -u frameflow-backup.service -u frameflow-restore-rehearsal.service --since today ``` The units intentionally do not delete old backups. Apply an operator-reviewed retention or immutable-storage lifecycle policy at the off-host destination, and keep at least one verified recovery point outside that policy's deletion window. Re-run verification after every off-host transfer and before any restore: ```bash ./scripts/verify-backup.sh /mnt/encrypted-backups/frameflow/frameflow-20260731T040000Z ``` Verification is read-only and does not connect to or modify the running stack. Do not proceed when it reports a checksum, manifest, dump-format, or archive-path failure. At least monthly, restore a recently transferred backup into the isolated rehearsal environment: ```bash FRAMEFLOW_RESTORE_METRICS_DIR=/var/lib/frameflow/metrics \ npm run restore:rehearse -- \ /mnt/encrypted-backups/frameflow/frameflow-20260731T040000Z ``` The rehearsal verifies the backup before making any Docker call. It creates a randomly named internal network, three new empty volumes, and temporary PostgreSQL, Redis, and MinIO containers without publishing host ports. Every resource carries both a restore-rehearsal marker and a run-specific label. Cleanup rechecks both labels and refuses to remove a resource if either label differs, so the script never removes the live Compose project or an unrelated similarly named resource. PostgreSQL validation restores the custom-format dump with `--exit-on-error`, checks migration history and core tables, and queries representative business tables. Redis validation loads the restored persistence, checks `PING`, `DBSIZE`, loading state, and AOF health. MinIO validation starts the production-pinned server image and uses the pinned `mc` client to check administrative readiness and list restored buckets and objects. On success, the script cleans the temporary resources and atomically publishes timestamp, duration, PostgreSQL table-count, Redis key-count, and MinIO object-count metrics for Node Exporter. Failed rehearsals are cleaned automatically. For incident diagnosis only, add `--keep-on-failure`; the command prints the exact retained resource names and label values. Inspect those labels before manually removing anything. The `FrameFlowRestoreRehearsalStale` alert fires when no successful rehearsal metric exists or the last success is older than 35 days. Verify both disaster-recovery freshness metrics through Prometheus: ```bash curl -fsS --get http://127.0.0.1:9090/api/v1/query \ --data-urlencode 'query={__name__=~"frameflow_(backup|restore_rehearsal)_last_success_timestamp_seconds"}' ``` For recovery, stop application writers first: ```bash docker compose --env-file .env.production -f docker-compose.prod.yml stop web api worker ``` Restore `postgres.dump` with `pg_restore` and unpack the Redis/MinIO archives only into newly created, empty recovery volumes. Restoring volume archives over live data can leave stale objects and is not supported. Start dependencies, run the `migrate` service, then start API, worker, and web and complete the media/model checks above. ## Upgrades Use an immutable `FRAMEFLOW_VERSION`, build the new images, take a backup, and then recreate services. The migration job completes before new API and worker containers become healthy. ```bash ./scripts/backup.sh /mnt/encrypted-backups/frameflow docker compose --env-file .env.production -f docker-compose.prod.yml build docker compose --env-file .env.production -f docker-compose.prod.yml up -d docker compose --env-file .env.production -f docker-compose.prod.yml ps ``` Rollback application images only after confirming that the new migration remains backward compatible. Restore the database from backup when a migration is not reversible.