EP12: Benchmarking low-latency live streaming and the current baseline¶
| Recap | EP11 "What Netflix's seamless switching teaches us" |
|---|---|
| Next | EP13 "P2P direct connect: faster than CDN for the last mile" (entering Module 3) |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Recall the repeatable current baseline: P2P direct connect runs at ~70ms; when distributed through Edge, WHIP ingest runs at ~108ms and SRT ingest at ~386ms.
- Design a same-caliber, reproducible low-latency benchmark, not conflating vendor marketing, a single observation and peer measurements.
- Understand the current evidence boundary: without peer-condition data for Huawei Cloud or Tencent Cloud, this episode makes no superiority claims.
1. Opening hook (script notes)¶
A comparison table is easy to make; a credible comparison is hard. If the device, encoder, network, timestamp caliber or sample window differ, the numbers can't be compared directly. This episode first fixes this project's measured baseline, then gives a peer-benchmark protocol; until competitor data is filled in, don't dress speculation up as a conclusion.
2. The current confirmable baseline¶
Under the current test conditions, the end-to-end latencies actually proven and recorded are (data from
a single-machine, single-stream, time-boxed Live Streaming Latency Performance Test Report, last
updated 2026-09-29):
| Path | Measured | Interpretation limit |
|---|---|---|
| P2P direct (bypassing Origin/Edge) | ~70ms | Proven under specific experimental conditions; not a commitment for all NAT, regions and networks |
| Edge distribution, WHIP ingest | ~108ms | WHIP is WebRTC-based; ingest has no fixed receive window; must be read together with test device and network conditions |
| Edge distribution, SRT ingest | ~386ms | About 278ms higher than WHIP, mainly from the SRT receive window (TSBPD) at ingest — a fixed delivery delay deliberately enlarged for retransmission headroom; weak networks add retransmission buffering on top |
These are measurement results, not an SLA, and not fixed constants for every environment. The gap among the three is mainly determined by ingest protocol structure, not codec: the same test used the same encoding config throughout, and the gap between the WHIP and SRT ingest paths comes from the SRT receive window — a fixed delay deliberately enlarged to give room for retransmission on weak networks. The encoding format used for the push stream (H.264/HEVC) is currently the basis the player uses to pick a route according to decode capability — HEVC always goes through Edge and doesn't participate in P2P direct connect, while H.264 can attempt P2P or fall back to Edge — that's a routing rule, not an independent latency variable that's been isolated by a controlled experiment; without an experiment that separates out the codec's own effect, codec differences shouldn't be folded into this set of latency numbers.
For reference, this project's internal SLA target for the P2P direct-connect path is P80 ≤300ms and P99 ≤2s; the current measured value of ~70ms is clearly better than that target, but the sample size and scenario coverage are still limited, so it shouldn't be generalized to every network environment.
End-to-end latency follows EP08's caliber: correlate the same frame's capture time with its actual render time, and disclose clock-calibration error and invalid-sample rate. A standalone ICE RTT, RTCP RTT, packet-receive instant or decode-complete instant cannot replace this metric.
3. Why no competitor superiority conclusion is possible yet¶
There is currently no raw data for Huawei Cloud or Tencent Cloud collected under peer conditions with this project: same capture source, devices, encoding parameters, network path, test window, play protocol, timestamp method and statistical caliber. So this episode can't answer "who's faster" or "who's more stable".
Vendor capability pages and nominal latency can be used to define the configuration to test, but are not local peer measurements. For example, Tencent Cloud's Low-latency Live Broadcasting (LEB) marketing page describes latency with a range-style claim — "millisecond-level / within 1 second" — which is itself not a single measured data point; subtracting or ranking that kind of public claim directly against this project's single measured number is already an apples-to-oranges comparison before you even ask who's faster — you'd be comparing two numbers of a fundamentally different nature. Architecture analysis can raise hypotheses — e.g. P2P may reduce forwarding hops, HEVC may lower bandwidth needs at equal subjective quality — but these hypotheses can't replace measurement, let alone justify claims about Huawei Cloud's or Tencent Cloud's relative merit.
4. A protocol for a credible benchmark¶
4.1 Fix the control variables¶
- Same capture source, resolution, frame rate, rate-control mode, GOP and audio config.
- Same batch of play devices, browser/native player versions and hardware-decode settings.
- Repeat tests in a similar time window, on the same access network and region, and record the network type (wired, WiFi, 4G/5G).
- Compare the same codec and protocol combination separately. If a product doesn't support the same combination, state it explicitly as a capability difference — don't attribute the latency of different combinations directly to the vendor. This project's own HEVC push streams currently only go through Edge and don't participate in P2P direct connect (see §2) — a real example of why stats must be broken out by codec+protocol combination rather than lumped together.
4.2 Fix the measurement caliber¶
- End-to-end latency uses frame-level capture time to actual render time, recording clock offset, uncertainty and invalid-sample rate.
- First-frame uses a unified "valid play request accepted" to "first non-placeholder frame actually rendered" caliber.
- Rebuffering uses EP08's unified event threshold, exclusions and effective-watch-time denominator.
- Record ICE RTT, RTCP RTT and end-to-end latency separately, not merged into one "network latency".
4.3 Look at distribution and failures, not the single best run¶
Each test unit should report at least sample count, success rate, P50, P80, P95, P99, maximum and confidence interval, and separately list connection failures, timeouts and invalid samples. The P99 method and minimum sample size must be fixed in advance, so no one picks a favorable statistic after the fact.
4.4 Make it reproducible¶
Publish the test date, region, devices, versions, config, scripts, raw anonymized data and exclusion rules. Re-test after a product version or key config changes, so old results don't long represent the current product.
5. How the results table should be written¶
Before peer data exists, the results table records only the evidence state, not speculative values:
| Subject | Peer-benchmark status | Current conclusion |
|---|---|---|
| This project | Current baseline available | P2P direct ~70ms; under Edge distribution, WHIP ingest ~108ms and SRT ingest ~386ms — scope limited to the recorded test conditions |
| Huawei Cloud | No peer benchmark this round | No latency/stability superiority judgment |
| Tencent Cloud | No peer benchmark this round | No latency/stability superiority judgment |
For later re-tests, add same-caliber distributions and raw-data links directly, rather than writing a conclusion first and hunting for supporting numbers.
6. Diagrams¶
6.1 From benchmark protocol to conclusion¶
Same capture source, devices, encoding and network conditions
│
▼
Unified frame-level end-to-end latency / first-frame / rebuffer definitions
│
▼
Multiple rounds, recording success, failure and invalid samples
│
▼
Report P50/P80/P95/P99, sample size and confidence intervals
│
▼
Publish config, scripts, raw data and exclusion rules
│
▼
Only same-caliber data yields a product comparison conclusion
6.2 The current evidence boundary¶
Current baseline: P2P direct≈70ms / Edge-WHIP ingest≈108ms / Edge-SRT ingest≈386ms
│
├── usable as the baseline for later regression and peer tests
│
└── cannot be extrapolated directly into an SLA for all environments
Huawei Cloud / Tencent Cloud: no peer benchmark this round
│
└── no superiority conclusion; no substituting vendor marketing
7. Wrap-up and next episode (script notes)¶
The current facts are just three baseline numbers: P2P direct connect at ~70ms, and under Edge distribution, WHIP ingest at ~108ms and SRT ingest at ~386ms — the gap comes mainly from ingest protocol structure, not codec. Without peer competitor measurements, make no Huawei Cloud / Tencent Cloud superiority judgment. What's truly reusable is the benchmark protocol: control variables, unified frame-level caliber, report the full distribution and failures, publish raw data.
That completes Module 2. Next episode enters Module 3: how P2P direct connect is established via ICE checks, and how the product trades off among direct, Edge and TURN.