Adds an opt-in RedisIoAdapter (wired in main.ts when REDIS_URL is set) so socket.io room emits fan out across instances via Redis pub/sub — a client on replica B now receives messages emitted by replica A. Without REDIS_URL the in-memory adapter is kept (single-instance dev unchanged). Adds redis to docker-compose, REDIS_URL to the env contract, and smoke-realtime-cluster.mjs which proves cross-instance delivery fails without Redis and passes with it. Delivery is at-least-once at N>1; clients dedupe by message id (documented). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
5.5 KiB
Deploying IIOS
IIOS deploys as one runtime service (@insignia/iios-service) plus one datastore
(Postgres). Everything else in the repo is a library (SDKs) or a dev demo, not a
deployed service.
| Artifact | Type | How it ships |
|---|---|---|
@insignia/iios-service |
Runtime service | Docker image (this guide) |
iios-contracts, iios-adapter-sdk, iios-kernel-client, iios-testkit |
Library | published to npm; imported by host apps |
iios-message-web, iios-inbox-web, iios-support-web, iios-ai-web, iios-meeting-web, iios-community-web |
Web SDK (library) | published to npm; bundled into host-app frontends (optionally one CDN widget) |
apps/* (message-demo, ai-studio, route-admin…) |
Dev demo | not deployed |
The service is a modular monolith: a single NestJS process serves the HTTP API, the WebSocket (socket.io) gateway, the outbox relay, and the projectors together.
Build the image
Build context is the monorepo root (the image needs the workspace + lockfile):
docker build -f packages/iios-service/Dockerfile -t iios-service:latest .
The multi-stage build compiles the service + its workspace deps, generates the Prisma
client, and prunes to a self-contained prod bundle (pnpm deploy --legacy). The runtime
image runs as the non-root node user and starts via docker-entrypoint.sh, which runs
prisma migrate deploy then node dist/main.js.
Run it
docker run --rm -p 3200:3200 \
-e DATABASE_URL='postgresql://USER:PASS@HOST:5432/iios?schema=public' \
-e APP_SECRETS='{"portal-demo":"<secret>"}' \
-e IIOS_DEV_TOKENS=0 \
iios-service:latest
Health check: GET /health → 200. Metrics/ops: GET /metrics (relay lag, projection
cursors, retention snapshot counts).
Configuration & secrets contract
Every knob is an environment variable — see packages/iios-service/.env.example
for the full, commented list. Highlights:
- Secrets (inject from a vault, never bake into the image):
DATABASE_URL,APP_SECRETS(per-app JWT signing keys),ADAPTER_SECRETS(webhook HMAC keys). - ⚠️
IIOS_DEV_TOKENSMUST be0/unset in production. It exposes/v1/dev/*(unauthenticated token minting, webhook injection, chaos, retention sweep). This is the single most important prod-hardening flag. IIOS_CELL_IDtags which physical cell this instance serves — the hook for splitting noisy/regulated tenants into isolated cells later, with no code change.- Worker timers (
IIOS_RELAY_INTERVAL_MS,IIOS_RETENTION_SWEEP_INTERVAL_MS) — see scaling notes below.
Topology
host apps ──HTTPS/WSS──► [ LB / ingress ] ──► iios-service ×N ──► Postgres (pgvector)
(embed SDKs) (WS + sticky) (API+relay+ │
projectors) (+ Redis, + OPA/CMP/MDM
as external ports — later)
- Postgres — managed (RDS / Cloud SQL / Neon). One logical DB, tenant-isolated by scope.
- Redis — not required for a single instance. Set
REDIS_URLwhen you run multiple replicas with live chat: the built-in socket.io Redis adapter fans room emits across all instances, so a client on replica B sees messages emitted by replica A. WithoutREDIS_URLthe in-memory adapter is used (single instance). Verified byscripts/smoke-realtime-cluster.mjs(two instances + one Redis; cross-instance delivery fails without Redis, passes with it).Delivery is at-least-once across instances — with N>1 the inline send-emit and the relay's outbox-bridge emit can both fire for the same message, so clients should dedupe by message
id(the payload always carries one). - Platform ports (OPA policy, CMP consent, MDM, CRRE, SAS) — today in-process permissive
stubs (
LocalDevPorts). For production, point these at real external services; the service already calls them fail-closed.
Scaling & release strategy
Horizontal scaling is safe because the event core was hardened for it:
- The outbox relay claims rows with
FOR UPDATE SKIP LOCKED→ each event is relayed by exactly one replica; its co-located projectors process it once (idempotentclaim()+ the projection-cursor + idempotency-command ledgers give exactly-once effects). - The DLQ + replay path means a bad deploy loses nothing — fix and replay.
Rolling / blue-green deploys:
- Migrations: run
prisma migrate deployas a one-off pre-deploy job, then setIIOS_SKIP_MIGRATE=1on the replicas so N pods don't race the migration. (For a single instance, the default boot-time migration is fine.) - Use expand-contract (backward-compatible) schema changes so old and new pods can run against the same schema during the roll — this is the prerequisite for zero-downtime.
- Roll replicas with a
/healthreadiness gate; drain WebSocket connections onSIGTERM. - WebSocket ingress needs sticky sessions (and the Redis socket.io adapter once N>1).
Not yet built (prod-readiness gaps)
- CI/CD pipeline to build/push this image and run migrations.
- Real platform-port services (OPA/CMP/MDM) — today permissive stubs.
- Secrets manager integration (currently plain env).
- A schema-compatibility gate in CI to enforce expand-contract migrations.