EP07: How to strengthen resilience on weak networks¶
| Recap | EP06 "Publish/play security: signing, tokens and anti-hotlinking" |
|---|---|
| Next | EP08 "Video instrumentation: measuring real end-to-end latency" |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Explain what SRT / RTP native statistics each represent, and how a product combines multiple signals into degradation conditions — different protocols' counters can't be forced into one "unrecoverable loss rate".
- Understand that "weak-network resilience" is layered engineering, where detection and degradation are only two of the layers.
- Understand two easily-missed mechanisms that strongly affect real weak-network experience: preemption protection between P2P and the main publish, and how "reconnect" itself must be designed so it doesn't crush the system the instant the network recovers.
1. Opening hook (script notes)¶
"Degrade when there's loss" sounds simple, but in practice you first answer: does this counter count sequence-number gaps, retransmission requests, retransmitted packets, or media loss after the recovery deadline? SRT and RTP native stats differ in caliber; a product must interpret them separately, then combine loss, RTT, send-queue and available bandwidth into configurable action conditions.
2. Layer 1: interpret protocol counters first, then define product metrics¶
2.1 SRT and RTP native stats cannot be trivially unified into a ULR¶
The SRT library usually exposes loss, retransmission, dropped, NAK and send-buffer stats, but whether each field is sender- or receiver-side, a packet event or a unique packet, and whether it includes duplicate retransmissions all depend on the SRT version's API. You cannot derive the unrecovered packet count with "cumulative loss minus cumulative retransmit": one loss can trigger several retransmissions, and the counters may not even share a statistical domain.
RTP/RTCP's cumulative packets lost and fraction lost are computed mainly from expected vs
received RTP sequence numbers. They reflect a receive gap but don't inherently tell the app whether
the gap was later recovered by RTX/FEC in time for the decode deadline. NACK count, RTX received
packets, pre-decode drops and stutter are all signals at different layers.
So the product layer should keep the native metric names and units and define its own business metric separately — e.g. "proportion of media units still missing after the recovery/playout deadline". Only when the receiver can correlate original-packet identity, RTX/FEC recovery outcome and the playout deadline can it compute such a derived metric; don't just rename an arbitrary raw counter as ULR.
2.2 Thresholds and sampling are product parameters, not protocol constants¶
Sampling window, degrade threshold, recover threshold, consecutive-hit count and cooldown should all be product parameters configured per protocol, direction, content layer and terminal capability, not constants from SRT or RTP standards, and you shouldn't assume the two protocols share one set. Production values need calibration from load tests and real network data, recorded per version.
Design usually uses a rolling window, consecutive hits and hysteresis: the degrade threshold is higher than the recover threshold, so degrading can be fast and recovering more cautious. This filters single spikes and avoids oscillating up/down near a network's critical point. Training diagrams should show parameter names only, not turn one experiment's values into universal thresholds.
2.3 Separate upstream publish degradation from downstream playback ABR¶
- Upstream publish degradation happens on the encoder→Origin link. The sender can lower target bitrate, frame rate or resolution, pause upper layers, and renegotiate if needed; whether a given action is seamless, needs a keyframe, or rebuilds the session depends on the encoder and protocol implementation — you can't promise "zero cost" or a fixed interruption length uniformly.
- Downstream playback ABR happens on the Origin/Edge/P2P→viewer link. The player picks an existing layer based on that viewer's goodput, RTT, loss, receive/playout buffer and decode state, and must not change global publish quality because one viewer is on a weak network.
The two control loops can share the observation system but differ in target, threshold and blast radius. Upstream anomalies affect all viewers, so actions should be more conservative; downstream ABR decides per session and usually prefers switching to a lower existing layer.
3. Layer 2: Simulcast multiple layers — which layer the detection result acts on¶
EP05 covered Simulcast's layering. Here we add the control domain: the publisher can adjust the layers it sends based on the overall uplink budget; the player selects already-published layers per session. Switching should pick a point the receiver can independently decode, coordinating keyframe requests, dependencies and buffer state. The detection module outputs network state with confidence, and the policy module decides which link, which layer, and when to switch.
4. Layer 3: main publish first, P2P is the first thing sacrificed¶
A publisher's uplink bandwidth is finite; once it carries both "the main publish to Origin" and "P2P direct to multiple viewers", the two contend on a weak network. The principle is direct: the main publish quality is always the top priority, and P2P is an add-on that can be sacrificed at any time. This lands as three hard constraints:
- Independent send queues and bandwidth budgets: the main publish (Origin WHIP) and every P2P output keep their own send queue, congestion-control state and pacing, physically isolated; P2P's send logic never steals the main publish's scheduling slot.
- Absolute priority, not "fair sharing": the main publish's queue and bandwidth budget always come before P2P — not "split the bandwidth", but "the main publish eats first, P2P uses what's left".
- Active monitoring + active removal: the publisher continuously monitors local uplink conditions; once a congestion threshold is hit, it immediately terminates one or more P2P connections without waiting for any central command — and after recovery it does not automatically reclaim P2P; it must request a quota from the scheduler anew.
Why "not waiting for a central command" matters: if it had to report to the control plane and wait
for a "disconnect P2P" command, that round-trip is itself an indeterminate delay on a weak network.
P2P removal is designed so the publisher can decide and execute locally and immediately, with no
network round-trip — this is also verifiable in the code: the actual removal is a direct local call
to signaling (SendClose), not a report-and-wait-for-reply round trip.
Current implementation: only one signal — RTT — is actually wired up
The architecture design doc envisioned a monitoring-signal set that includes RTT, loss rate,
send-queue length and available bitrate (docs/design/ppcdn-architecture-design.zh-CN.md
§6.4), but the decision logic ppobs actually ships
(plugins/obs-webrtc/uplink-qos-policy.cpp) only wires up one of them: uplink RTT exceeding a
threshold is treated as congestion, with a default threshold of maxRttMs=300ms, checked
every 2 seconds by WHIPOutput::CheckUplinkQos. By default it removes the 1
most-recently-established P2P connection (maxPeersToDropAtOnce=1, configurable to remove
all of them), then enters a 10-second cooldown (cooldownMs=10000) before re-evaluating.
Loss rate, send-queue length and available bitrate are not yet wired into the decision
condition — that's a known gap between design intent and the current implementation, and it
shouldn't be dressed up as "all four signals are already combined into the decision."
5. Layer 4: reconnect must be designed too, or the recovery moment becomes the second outage¶
After a network blip the client tries to reconnect — but if "how to reconnect" is poorly designed, it creates new problems the instant the network recovers.
5.1 Backoff strategy: the control connection already has jitter, media reconnects are still fixed-interval¶
The general principle behind "should we add backoff + jitter" is: exponential backoff solves "don't keep blindly retrying the instant something fails" — avoiding hammering an already-unstable link before the network has really recovered; jitter solves a different problem — if thousands of clients that dropped at the same moment all retry at exactly the same points on the same curve, then the instant the network recovers, every one of those requests slams in together. That's the "thundering herd" effect: the recovery moment itself creates a traffic spike. A little random jitter staggers those retries in time.
But in the current code, only part of these two principles is actually in place — it would be wrong to say broadly that "every reconnect does jittered exponential backoff":
- The node↔control-plane control connection (ppmmx's heartbeat/control WebSocket to ppcenter,
ppmmx/internal/mmxcontrol/client.go) does implement full bounded exponential backoff + jitter: an initial 1-second wait, doubling after every failure, capped at 30 seconds, with random jitter of up to half the current backoff duration layered on top. - Media-path reconnects are currently all fixed-interval, with no exponential growth and no
jitter: Edge retries a failed pull from Origin after a fixed 5 seconds (
retryPauseinppmmx/internal/staticsources/handler.go, the same constant shared by every static-source type); the player-side WHEP reader waits a fixed 2 seconds (MediaMTXWebRTCReader.retryPauseinpplayer/ppplayer.mjs), and the ABR control WebSocket waits a fixed 3 seconds (reconnectDelayin the same file).
That doesn't mean the thundering-herd risk is gone on these paths — if an Edge node drops at scale, every player attached to it really will try to reconnect at roughly the same moment — it's just that this part of the reconnect logic hasn't adopted jitter yet, a known and still-open point for future optimization. The diagram below (6.3) illustrates the general "jittered exponential backoff" mechanism itself, which today is fully in effect only on the control connection.
5.2 Nodes carry it themselves: reconnect even when the control plane is offline¶
An Edge node caches the Origin address template and credentials for the current stream. If the
Edge→Origin link blips — even if the control plane (ppcenter) itself is unavailable at that
moment — Edge doesn't need to wait for it to recover: reconnecting just means locally re-resolving
the cached source config (the staticsources package's resolveSource, a pure string substitution
that sends no request at all), then connecting straight to Origin with the cached credential. So
whether the control plane is online or offline has no bearing on this particular reconnect. Per the
previous section's finding, this reconnect loop currently runs on a fixed 5-second interval, not
exponential backoff. Only when reconnects keep failing, the Origin is confirmed to be genuinely
down, or an upper layer decides to terminate does it actually give up and report upward.
This means "a brief control-plane failure" and "a media-path outage" are fully decoupled: as long as the cached credential hasn't expired and the Origin is alive, whether the control plane is online during that window may be entirely imperceptible to viewers.
5.3 Honest boundary: this is not "seamless failover"¶
If the Origin truly fails (not a network blip but the node going down), the publisher must republish to a new Origin before playback requests can use the new path — cross-Origin seamless failover is not promised at this stage, and the switch has a perceptible interruption. Worth stating clearly, to avoid exaggerating it into "any failure recovers invisibly".
6. Diagrams¶
6.1 Per-protocol observation and per-direction control¶
SRT native stats ──┐
├──▶ interpret per protocol + rolling window / consecutive hits / hysteresis (all product params)
RTP/RTCP/RTX ──────┘ │
▼
joint decision on RTT / queue / goodput / buffer
│
┌──────────────────┴──────────────────┐
▼ ▼
upstream publish control downstream session ABR
adjust bitrate/fps/publish layers pick an existing layer per viewer
affects the whole stream, more doesn't change others' publish quality
conservative policy
6.2 Congestion detection → active P2P removal → re-request after recovery¶
Publisher checks every 2 seconds: local uplink RTT (the only signal currently wired up;
design intent also includes loss rate / queue length / available bitrate)
│
▼
RTT ≥ threshold (default 300ms)?
│
┌────┴────┐
no yes
│ │
continue ▼
normally immediately terminate 1 or more P2P connections (local decision, no central command)
│
▼
main publish bandwidth budget restored
│
▼
report to center: P2P quota released; enters a 10-second cooldown
│
▼
after network recovers, no automatic reclaim — must request a new P2P quota from the scheduler
6.3 Jittered backoff: from thundering herd to staggered recovery (currently fully in effect only on the node↔control-plane connection — see §5.1)¶
[No jitter, one backoff curve]
Failure
│
▼ (all clients use the same backoff times)
Retry 1: all clients fire at t=1s
Retry 2: all clients fire at t=2s
│
▼
The instant the network recovers, all retry together → spike, may knock the service down again
[Exponential backoff + random jitter]
Failure
│
▼ (each client adds a random offset within the window)
Retry 1: clients spread across t=0.8s ~ 1.2s
Retry 2: clients spread across t=1.6s ~ 2.4s
│
▼
Retries are staggered in time → smooth recovery, no concentrated shock
6.4 Edge autonomous reconnect: what to do when the control plane is offline¶
This path is confirmed implemented (docs/ppcdn-development-progress.zh-CN.md §3.5, status
"completed", which added the POST /internal/mmx/v1/origin/resolve endpoint) — it isn't just
sitting in the architecture design doc's list of MUST requirements.
Edge ←→ Origin link blips
│
▼
Edge dynamically resolves the Origin address + a short-lived media credential from ppcenter;
the result is cached in memory
when ppcenter is unavailable → reuse the unexpired cache to keep reconnecting
(once the credential expires, it may no longer be used)
│
┌────┴────┐
connects OK keeps failing / Origin confirmed dead / cache expired
│ │
▼ ▼
largely imperceptible terminate the reader, report upstream
to viewers
(fixed 5-second retry
interval; the node's
control-plane credential
and the Origin media
credential are independent
of each other)
7. Wrap-up and next episode (script notes)¶
This episode drew the statistical boundaries: SRT and RTP native counters can't be unified into a ULR via a single subtraction; sampling windows and up/down thresholds are product parameters that need calibration from measurement and version management. Control must also separate whole-stream upstream publish degradation from per-viewer downstream ABR. Add the main-publish preemption of P2P (currently wired up with only the RTT signal), jittered backoff on the control connection together with autonomous node reconnect (media-path reconnects are still fixed-interval retries today), and that's the weak-network resilience design of the current version — one that deliberately marks which parts are still design intent and which are already borne out in the code.
Next episode, a different angle — not how to cope with weak networks but how to measure "how much latency there actually is": video instrumentation and end-to-end latency measurement.