RustPBX as a Cluster: One Voice Plane Across Many Nodes
Single-box PBX is easy. The hard part starts when you run two nodes and discover that “shared database” doesn’t make phones work together: a call lives on exactly one node, registrations land wherever the load balancer sends them, and the console you’re logged into may have no idea that the extension you want to transfer is being served by its neighbor.
RustPBX 0.5 closes that gap with a distributed session registry. Here’s what it changes in practice.
The Three Problems of Naive Clustering
- Calls are stranded. Node A owns the call; the supervisor’s console is on node B. Clicking “hang up” does nothing, because B has no idea the call exists.
- Context evaporates. Your CRM attaches a ticket id to the session. The call transfers to a leg on another node — and the ticket id is gone.
- Failures leave ghosts. A node dies; its calls vanish from the active-call list without ever being reported as ended.
What the Session Registry Does

Owner-routed operations. Any node can answer, transfer, hang up, or set variables on a call it doesn’t host — the request is forwarded to the owning node automatically. One load balancer, three nodes, and a single control plane from any of them.
Session user_data replication. Application data attached to a session — CRM ids, campaign tags, agent context — replicates to peers, and transfer sub-sessions inherit it. Business context survives node handoffs.
# works from any node — routed to the owner transparently
curl -X PUT https://node-b/api/calls/active/<session_id>/userdata \
-H "Authorization: Bearer <token>" \
-d '{"customerId": "C-8192", "tier": "gold"}'
Heartbeat gauges. Each node reports liveness of the sessions it owns; pruned bindings are reported as offline rather than silently disappearing.
Cluster AMI Surface
The cluster endpoints make multi-node operations first-class:
| Endpoint | Use |
|---|---|
GET /cluster/list_calls | Every active call across the cluster |
GET /cluster/session_owner/{call_id} | Which node hosts a call |
POST /cluster/session_op | Owner-routed RWI command envelope |
POST /cluster/set_userdata / get_userdata | Cross-node session data |
GET /cluster/logs/recent · /follow | Peer logs from one place |
GET /cluster/evaluate_route | Route evaluation across the cluster |
Rolling Restarts That Nobody Notices
SIGTERM/SIGINT (or POST /shutdown) triggers drain-then-exit: the node stops accepting new calls, lets in-flight ones finish, then exits. Restart nodes one at a time and the registry routes around the draining node. For signaling-heavy environments, pair it with SipFlow cluster mode so capture continues on peers.
The Deployment Shape
- Shared MySQL/PostgreSQL (SQLite is single-node only)
- Identical config via GitOps; peer list in
[cluster].peers - Load balancer with source-IP or Call-ID affinity for SIP
- Config, routes, and IVRs on every node — the registry moves calls, not files
Full topology and step-by-step: Cluster Deployment.