LiveBoard / docsProject documentation

Reference

Distributed Runtime

Look up syntax, contracts, layouts, algorithms, and exact behavior.

LiveBoard can run several FastAPI backend containers behind a backend proxy. This required separating local socket ownership from room-wide collaboration semantics. Each replica owns only the WebSocket objects connected to that process. Redis carries room events, presence records, invalidation messages, and rate-limit counters across replicas.

Without Redis, a single process can broadcast to the sockets it owns, but two backend processes would become two isolated rooms. A user connected to replica A would not automatically reach a user connected to replica B. Redis closes that gap by acting as the shared notification layer between replicas while leaving each process responsible for its own in-memory WebSocket objects.

Client AReplica APostgreSQLRedis Pub/SubReplica BClient BClient Ccommit firstpublish after commit
Cross-replica fanout. A room is logical, not process-local. A replica commits durable operations to PostgreSQL, then uses Redis Pub/Sub so every other replica can forward the event to its local sockets.
Durable=PostgreSQL,Ephemeral=Redis∪BrowserDurable = PostgreSQL,\quad Ephemeral = Redis \cup Browser
Authority split. Redis is deliberately not a cache for canvas state. It coordinates events that are either transient or recoverable from PostgreSQL.
liveboard:canvas:{canvas_id}:events
liveboard:presence:{canvas_id}:connections
liveboard:presence:conn:{connection_id}
rate:{scope}:{bucket}
Redis keys. The scaled runtime uses Pub/Sub channels for canvas events, TTL records for presence, and fixed-window counters for shared rate limits.

Presence is user-level even when the same user has multiple tabs open. Redis stores connection ids with short TTLs. A join is broadcast when a user moves from zero connections to one or more, and a leave is delayed briefly so a refresh or tab handoff does not flicker the collaborator list.

The TTL detail is important in a distributed WebSocket system. If a backend process exits without running its disconnect cleanup, Redis will eventually expire its connection records. That means the presence layer is eventually self-healing instead of depending on perfect shutdown behavior. The short delayed leave avoids the opposite problem: making ordinary reconnects look like people rapidly leaving and rejoining.

socket accepted
  -> add connection id to canvas set
  -> write connection record with TTL
  -> if previous user connection count was 0: presence_join

socket closed
  -> remove connection id
  -> wait briefly
  -> if user connection count is still 0: presence_leave
Presence transition. The collaborator list is derived from Redis presence in scaled mode and from local sockets in single-process fallback mode.

Rate limits are shared across replicas when Redis is enabled. Authentication routes, HTTP API routes, cursor messages, preview messages, history messages, and durable writes each have separate counters. The write limit is the most important collaboration guard because it protects revision churn and history growth. Cursor and preview limits are intentionally higher so normal collaboration remains fluid.

The rate-limit behavior is tied back into collaboration recovery. When a durable write is rejected, the server does not let the rejected operation leak to peers as if it had succeeded. The writer receives the saved canvas snapshot and briefly stops interacting with the canvas while it syncs. Peers receive a preview reset so any transient movement they saw from that writer is discarded. The failure mode is visible instead of being hidden behind a best-effort patch over a rejected write.

docker compose up --build --scale server=3
curl http://localhost:3001/health
# {"ok": true, "postgres": true, "redis": true}
Scaled local runtime. The Docker Compose path runs backend replicas behind Caddy, with all replicas sharing PostgreSQL and Redis.

The single-process development path still works when REDIS_URL is unset: fanout, presence, and rate limits fall back to in-memory process state. That fallback supports local backend work, but it is intentionally not treated as a distributed mode.