Skip to content
kazma.
ع Star 6 Get Started

Multi-Region & HA

Kazma supports horizontal replicas when all shared state lives in Postgres.

┌─────────────┐
Users ────────►│ Load Balancer│ health: GET /health/ready
└──────┬──────┘
┌─────────────┼─────────────┐
▼ ▼ ▼
Kazma-1 Kazma-2 Kazma-N
│ │ │
└─────────────┼─────────────┘
▼
Managed Postgres (HA)
(settings, sessions, tasks, checkpoints)

Requirements (same in every region / replica)

Section titled “Requirements (same in every region / replica)”
EnvSame across replicas?
KAZMA_DATABASE_URLYes (one primary DB cluster)
KAZMA_SECRETYes (or OIDC-only)
KAZMA_VAULT_KEYYes
KAZMA_PUBLIC_URLPublic HTTPS URL (for OIDC)
Local kazma-data SQLiteHA compose demo shares the kazma_data volume (memory/cron/vault/docs). SQLite WAL + busy_timeout=5000 is the concurrency story — not full HA. Prefer a single writer, or keep --scale at 2.
Terminal window
# .env must set KAZMA_SECRET and KAZMA_VAULT_KEY
docker compose -f docker-compose.ha.yml --profile nginx up -d --build --scale kazma=2
curl -s http://127.0.0.1:9090/health/ready | jq .
  1. Deploy one primary Postgres (multi-AZ) with automated backups + PITR.
  2. Deploy Kazma as a service in each region (or one region + CDN).
  3. Point LB health checks at /health/ready (503 = remove from pool).
  4. Use sticky sessions only if you keep ephemeral local state; with Postgres cutover stickiness is optional.
  5. Cross-region active-active with two write DBs is not supported — use one write primary (or a multi-primary DB product that you fully operate).
LayerAction
App replica diesLB stops routing after failed /health/ready
Postgres primary failsManaged failover (RDS/Cloud SQL); apps reconnect via pool
Region lossFail over DNS to secondary region apps still pointed at healthy PG

See also: SAAS_AND_POSTGRES.md, DISASTER_RECOVERY.md.