SipFlow Deep Dive: Signaling Capture That Scales to Clusters
When a call misbehaves, the CDR tells you what happened and the media counters tell you how it sounded. Neither tells you exactly what was on the wire. For that you need packet-level signaling — captured continuously, indexed by call, queryable weeks later, and cheap enough to leave on permanently.
That’s SipFlow.

Capture Without a Tap
SipFlow mirrors every SIP message and RTP packet from inside the proxy — no port mirroring, no external capture box, no PCAP files to fetch. Each packet is stored with:
- Microsecond timestamp
- Source/destination IP:port
- Message type (SIP vs RTP)
- Full payload
Packets land in a compact binary format (ZSTD-compressed) with the on-disk layout under your control:
[sipflow]
type = "local"
root = "/data/sipflow"
subdirs = "hourly" # none | daily | hourly
flush_count = 500
flush_interval_secs = 10
The Index: FlowDB
Raw packets are useless without fast lookup. SipFlow maintains a FlowDB LSM-tree index that maps call_id → packet offsets, so a query for “show me call X” touches the index, not the whole archive. The SQLite side carries call metadata and message references; the binary side carries payloads. Query by call id and get the whole ladder back.
Cluster Mode: Consistent Hashing
One capture node becomes a bottleneck, so SipFlow scales out. In cluster mode, capture traffic is routed by consistent hashing on call_id — every packet of a call goes to the same capture node, so ladders stay complete even though the proxies are many:
[sipflow.remote]
flush_interval_secs = 5
[[sipflow.remote.nodes]]
udp = "10.0.0.2:6060"
http = "http://10.0.0.2:6060"
Deploy the standalone sipflow binary (src/bin/sipflow) as the capture server. Because the hash is on the call id, adding a node rebalances a share of calls, not all of them.
Backpressure and Loss Reporting
Capture must never take down the call path. SipFlow uses bounded queues with backpressure and reports per-IP loss when it can’t keep up. That’s the operational contract: under extreme load you lose some capture and get told exactly whose traffic was dropped, rather than the process growing without limit or stalling calls.
Metrics distinguish the pipeline stages (sqlite_* alongside sipflow_*), so you can see whether a bottleneck is ingest, flush, or index — not just “capture is slow.”
Getting Audio Back
Signaling capture includes RTP, so a call’s media can be reconstructed into a WAV for playback — useful when a recording was disabled but the evidence is needed anyway. Media export can target local storage or object storage with presigned downloads.
Retention and Offload
- Subdirs by hour/day keep directory sizes manageable and make pruning a filesystem operation.
- Object storage upload offloads long-term retention; local storage keeps the hot window.
- Pair with the archive addon for CSV/archive exports of call records when compliance wants structured data, not packets.
When to Reach for SipFlow
| Question | Tool |
|---|---|
| Why did this call fail? | CDR + call_error |
| Was the audio bad, and where? | CDR media evidence (RTCP) |
| What exactly was signaled, and when? | SipFlow |
| Can I hear the audio without a recording? | SipFlow media replay |
| Prove a carrier ignored a header | SipFlow + header comparison |
Guides: the SipFlow addon and SipFlow at scale.