Postgres & SaaS Cutover
SaaS multi-user + full Postgres cutover
Section titled “SaaS multi-user + full Postgres cutover”What is on Postgres when KAZMA_DATABASE_URL is set
Section titled “What is on Postgres when KAZMA_DATABASE_URL is set”| Store | Table(s) | Module |
|---|---|---|
| Config / settings / platform users | kazma_settings | config_store.py |
| Chat sessions | kazma_chat_sessions | session_manager.py |
| Swarm tasks + metrics | kazma_swarm_tasks, kazma_swarm_worker_metrics | task_store.py |
| LangGraph checkpoints | LangGraph internal schema | AsyncPostgresSaver, opened by kazma_core/checkpoints_pg.py (agent and gateway). Pruned since 2026-09-27: each chat keeps its newest 200 checkpoints, one idle for checkpoints.retention_days (Settings → System → Chat step history, default 30; 0 keeps all) its newest 10 |
| Web sessions | still ConfigStore keys (also Postgres via settings) | web_sessions.py |
SQLite remains the default when no database URL is set (tests, local single-node).
Memory search (pgvector)
Section titled “Memory search (pgvector)”Setting KAZMA_DATABASE_URL also promotes V2 dense recall from sqlite-vec
to pgvector in the same database (table kazma_memory — the
memory.backends.vector.collection setting — cosine, HNSW index), if that
Postgres can hold vectors. The cognitive store (memory_state.db) stays
SQLite until you set KAZMA_MEMORY_STATE_ROLE=primary.
Neither pgvector nor Qdrant is needed: without them memory’s meaning search is exact over every memory, locally (sqlite-vec or NumPy, about 1 ms per 1,000 memories).
The server must ship the extension. postgres:16-alpine — the image in
docker-compose.postgres.yml and docker-compose.ha.yml — does not; use
pgvector/pgvector:pg16 (or a managed Postgres that offers pgvector).
- Kazma creates the extension and table on first use when its role may.
Otherwise, as a superuser:
CREATE EXTENSION IF NOT EXISTS vector;in the Kazma database, and grant the roleCREATEon its schema. - Kazma checks at boot and on every memory search (cached one minute):
- pgvector picked automatically (no explicit choice) on a Postgres without the extension: one INFO line, memory stays on sqlite-vec, and Settings → Memory shows Vector: full (local) with the reason.
- pgvector chosen in Settings but unusable (no extension, role may not create it, table of another vector size, server down): one WARNING naming the fix, and the Settings banner says the same.
- Test vector in Settings → Memory runs the same check on demand. When pgvector was picked automatically and the extension is missing, it tests the store actually in use (local sqlite-vec) and reports Vector OK with a note saying why pgvector is not used; a pgvector you chose that cannot work still reports Vector failed with the fix.
- Rebuild embeddings once if you already have history: Settings → Memory → Rebuild embeddings (upserts into pgvector).
- Changing embedder size: the table is sized by the embedder and never
resized. Point
memory.backends.vector.collectionat a new name, then rebuild embeddings. - Kill-switch:
KAZMA_PGVECTOR=0(sqlite-vec on purpose, no check). Explicit Qdrant in Settings is never overridden.
Moving an existing postgres:16-alpine database to pgvector/pgvector:pg16:
dump and restore, do not just swap the image on the same volume. Alpine uses
musl and the pgvector image glibc; text indexes built under one collation are
out of order under the other. With Kazma stopped, dump the whole database
from the old container (pg_dump -Fc), restore it with pg_restore into a
pgvector container on a new volume, point KAZMA_DATABASE_URL at it and
start Kazma. Keep the old volume until the new one has run for a while.
(scripts/pg_backup.py dumps only Kazma’s own tables — right for a shared
database, not for moving a whole one.)
Postgres-primary recall (state.role=primary) is ILIKE + pgvector RRF,
not ILIKE-only.
Cutover procedure
Section titled “Cutover procedure”- Install extras:
pip install -e ".[postgres]" - Start Postgres:
Terminal window docker compose -f docker-compose.postgres.yml up -d db - Migrate all stores:
Terminal window export KAZMA_DATABASE_URL=postgresql://kazma:PASSWORD@localhost:5432/kazmapython scripts/migrate_sqlite_to_postgres.py --data-dir kazma-data - Run Kazma with the same URL (compose sets it automatically).
- Smoke:
- Login (user / secret / OIDC)
- Chat history across restart
- Swarm task list
- Settings persist
Multi-user UI
Section titled “Multi-user UI”/login— User · Secret · SSO- Settings → Account — users + tenants (admin)
- Header — role badge + logout
Create admin:
from kazma_core.security.platform_rbac import create_local_usercreate_local_user("admin", "long-password-here", role="admin")Env checklist
Section titled “Env checklist”KAZMA_DATABASE_URL=postgresql://…KAZMA_DB_BACKEND=postgres # optional forceKAZMA_PRODUCTION=1KAZMA_VAULT_KEY=…KAZMA_SECRET=…KAZMA_PUBLIC_URL=https://…# OIDC optionalKAZMA_OIDC_ISSUER=…KAZMA_OIDC_CLIENT_ID=…KAZMA_OIDC_CLIENT_SECRET=…Backup & disaster recovery
Section titled “Backup & disaster recovery”Kazma backs up its own Postgres tables automatically (added after the 2026-08-14 incident in which another app dropped Kazma’s tables from a shared database):
- Every 6 hours, automatic: the backup/export loop dumps exactly the
tables in
kazma_core.db.pg_backup.KAZMA_PG_TABLESto{kazma-data}/backups/pg/pg_shared_<epoch>.dump(custom-Fcformat, atomic write, magic-validated; 3 local dumps kept, restic keeps the history →backups.pg.retention/KAZMA_PG_BACKUP_RETENTION). First dump ~2 min after boot. The weekly deep drill restores the newest one into a scratch database, checks it and drops it (on by default since 2026-09-27; Disaster recovery). - Manual dump now:
python scripts/pg_backup.py backup - Restore:
python scripts/pg_backup.py restore --latest(--file <name>,--dry-run,list). Restores only Kazma’s own tables — a foreign app sharing the database is never touched. - Boot guard: if a required table is missing at boot, the server logs a
CRITICAL with the restore command instead of limping along with
UndefinedTableerrors. - Kill-switch:
KAZMA_PG_BACKUP_ENABLED=0(orbackups.pg.enabled=false). - The dump is table-filtered on purpose — never a whole-DB dump — so no foreign app’s data leaks into Kazma’s backups.
Use docs/ops/DISASTER_RECOVERY.md plus pg_dump for Postgres. After restore, secrets must match (KAZMA_VAULT_KEY, KAZMA_SECRET).