Skip to content

EP09: Latency stability matters more than average latency

Recap EP08 "Video instrumentation: measuring real end-to-end latency"
Next EP10 "How to quantify video quality"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Explain why "average latency" is a lying metric, and which statistics actually reflect real experience.
  2. Understand that several apparently latency-unrelated designs here are really paying for stability — the jitter buffer, P2P's connection budget and the SLA percentile thresholds are fundamentally the same thing.

1. Opening hook (script notes)

EP08 covered how to measure the latency number. This episode tackles an easier-to-miss question: once you have a pile of latency numbers, which statistic should you look at? Looking only at the average may give a conclusion completely at odds with what viewers actually experience.


2. Why the average lies

Take one live stream where latency is 100ms for 99% of the time and spikes to 5s for 1% due to a network jitter — the average is about 149ms, which looks like a good number. But what does the viewer experience? The viewer remembers that 1% where the picture froze, not how good the other 99% "only 100ms" was. The average dilutes one severe stall into a seemingly harmless small number, while the stall itself doesn't disappear at all.

This is a named problem in distributed systems and real-time media: tail latency — what really drives user experience and complaints is often not "how fast it usually is" but "how bad and how frequent the worst small fraction is". The average is precisely the statistic least sensitive to the tail, because extreme values get averaged away by many normal samples. That's why it lies.


3. Percentiles: what P50, P80, P95, P99 each say

The right approach is percentiles, not the average:

  • P50 (median): half the samples are faster, half slower. It is naturally robust to extreme values — even if 1% of samples spike to 10s, P50 is barely affected and represents the "typical" experience.
  • P80/P95/P99: roughly 80%/95%/99% of samples are at or below the value, i.e. about 20%/5%/1% are slower. They describe where the tail begins to appear, not the worst tail value; judging extreme experience should also look at P99.9, the maximum and the over-threshold ratio.

This isn't a hypothetical example — this set of thresholds and the judging order is the actual P2P latency SLA running in this project's production environment today (ppcenter/internal/ws/play_hub.go's p2pSLAP80TargetMs=300, p2pSLAP80MaxMs=600, p2pSLAP99WarnMs=1000, p2pSLAP99MaxMs=2000): P80 target ≤300ms, acceptable ceiling ≤600ms; P99 ceiling ≤2s, judged after aggregating path=p2p samples by calendar day in Beijing time. The states must be mutually exclusive, so the same batch of data can't hit both pass and warn: judge fail first (P80>600ms or P99>2s), then pass (P80≤300ms and P99≤1s), and everything else valid is uniformly warn. That judging order itself resolves a small trap in how the rule reads on paper: if pass were judged directly as "P80≤300 and P99≤2000", it would overlap with the "P99 between 1000~2000ms is warn" range — only the fail-first, then-pass, everything-else-warn order keeps the two apart; the three rules can't be run as three independent checks in parallel.

Even though the thresholds are real production values, they're still just a current business decision, not a protocol constant — change the business goal or the sample baseline and these numbers should be recalibrated. And this decision currently only feeds backend display and an optional alert (AppLatencyDaily.slaStatus); it doesn't trigger any compensation, and the public website's wording has recently been narrowed from an external-facing "P2P latency SLA" to the more cautious "video metric monitoring," to avoid it being read as a commitment. The aggregation period, minimum valid sample count and invalid-sample handling still need to be specified separately — don't assume that fixing the thresholds automatically covers those boundaries too.


4. The jitter buffer: the core way to trade "unstable" for "stable"

Network delay is never a constant — queueing, momentary congestion and routing jitter all make packet arrival times uneven, and that unevenness is "jitter". EP04 covered SRT's TSBPD receive window and WHIP's reorder buffer from the angle of "how to handle loss"; from another angle, these two mechanisms do the other half of the same thing: deliberately let data wait a little longer in a buffer, flattening uneven arrival into a fixed, predictable playout cadence.

The trade is direct: more wait budget absorbs more arrival jitter but adds playout latency. Modern real-time jitter buffers are usually not a fixed duration but flex with arrival jitter, loss/ retransmission, frame rate, decode cost and the latency target; they shrink when the network steadies and expand when jitter grows.

pplayer's own default behavior is a concrete example of this: if nobody explicitly specifies a buffer length, it does not set any fixed jitterBufferTarget/playoutDelayHint — it just lets the browser's own continuously adaptive jitter buffer take over. "Not specified" is treated here as "let the browser decide," not as quietly falling back to some default number. Only when a viewer explicitly asks for a fixed buffer length via ?bufferMs= does that value override the browser's adaptive behavior, and even then it's clamped to the [100, 1000] millisecond range (MIN_BUFFER_MS/MAX_BUFFER_MS in pplayer/buffer-config.mjs). In other words, "dynamic flexing" isn't an abstract description — it's the actual behavior of the default path itself; a fixed buffer is the exception that only happens once a viewer actively opts out of that dynamic flexing.

Evaluate jitter buffers not by average latency alone but together with target latency, actual buffer duration, late drops, stutter and recovery speed — don't misjudge "stable latency" as "the bigger the buffer the better".


5. The connection budget is "how long you're willing to wait before giving up and falling back," not a first-frame ceiling

Earlier versions ran P2P and Edge in parallel as a "dual-path race": both paths were established at the same time, whichever produced a frame first won, bounded by a 500ms window limiting "how long to wait for the candidate path." That parallel race has since been replaced with a sequential model: the play decision now branches upfront on the server (docs/design/ppcdn-architecture-design.zh-CN.md §6.2) — HEVC clients get edge-only right away and never attempt P2P at all; H264 clients that also pass NAT eligibility get p2p-connect and try P2P alone, with the Edge address kept only as a sequential fallback after failure, never established at the same time as P2P. pplayer's source still keeps the old "dual-path race" implementation function, but its comment plainly says it's "kept for rollback ... no longer wired to the play decision"; the server still sends raceWindowMs/connectTimeoutMs fields in its response (for backward compatibility with older clients), but the current client's comment notes they "described the old symmetric race and no longer apply" — the fields weren't removed, they're just no longer used to make path-selection decisions.

The time budget that actually governs behavior today comes in two different orders of magnitude: waiting for the publisher to answer the SDP offer is capped at 5 seconds (P2P_HANDSHAKE_TIMEOUT_MS, a pure signaling round-trip, unrelated to connection setup or decode); after the SDP exchange completes, there's a grace period of max(4000, connectTimeoutMs + 2000) ms (starting at 4 seconds by default) to wait for the first frame to actually decode, and only once that times out does it fall back to Edge (graceMs in pplayer/main.js). These numbers are an order of magnitude bigger than the old 500ms, but the lesson behind them is the same as before: this is the budget for "how long you're willing to wait on P2P before giving up and falling back," not a promise that the first frame will appear within that time — DNS, auth, ICE, the Edge connection, keyframe waiting, decode and render all still happen independently, outside this budget. Path-decision/fallback time and first-frame time should be observed separately, with their own SLA for first-frame — don't treat the former's number as a ceiling on the latter.


6. Diagrams

6.1 Same data, two stories from average and percentile

A set of latency samples (first-frame time over 100 plays):
98 at 100ms, 2 at 5000ms

[Average]
(98 × 100 + 2 × 5000) / 100 = 198ms
   → conclusion: "experience is pretty good"

[Percentiles]
P50 = 100ms   → typical experience is indeed fine
P99 = 5000ms  → the slowest tail is already in the 5-second range
   → conclusion: "most people are fine, but there's a clear long tail"

Same data: the average flattens the tail; percentiles put it on display

6.2 Jitter buffer: trade fixed latency for a stable cadence

[No buffer: uneven packet arrival]
pkt1(10ms) pkt2(80ms) pkt3(15ms) pkt4(120ms) pkt5(20ms) ...
   │
   ▼
Playout cadence speeds up and slows down; the picture alternates smooth and stuttery

[Adaptive buffer: play by the current jitter estimate]
pkt1 pkt2 pkt3 pkt4 pkt5 ... wait in the buffer for their "scheduled playout time"
   │
   ▼
Playout cadence steadies; the target buffer expands or shrinks with network state

Cost: some added latency       Benefit: fewer late drops and stutters

6.3 SLA decision: why P80/P99 buckets, not the average

Daily P2P latency samples
        │
        ▼
   compute P80 and P99 (not the average)
        │
judge fail first: P80>600ms or P99>2s
        │ no
        ▼
judge pass: P80≤300ms and P99≤1s
        │ no
        ▼
everything else is warn (states mutually exclusive)

(even with a low average, a poor P99 won't be judged pass)

7. Wrap-up and next episode (script notes)

The core of this episode is one line: stability cannot be answered by an average; look at clearly defined percentiles and tail samples. The jitter buffer should dynamically balance latency and stutter, P2P's connection budget is only "how long you're willing to wait before falling back," not a first-frame ceiling, and SLA states must be designed mutually exclusive with a clear sample caliber.

Next episode, from "is latency stable" to another topic equally misled by averages — how to quantify video quality.